Rethinking Molecular OOD Generalization via Target-Aware Source Selection

arXiv cs.LG Papers

Summary

This paper introduces SCOPE-Bench, a benchmark for evaluating molecular out-of-distribution generalization, and POMA, a framework using reinforcement learning to select source domains for domain adaptation, achieving significant error reductions on 3D molecular models.

arXiv:2605.13932v1 Announce Type: new Abstract: Robust prediction of molecular properties under extreme out-of-distribution (OOD) scenarios is a pivotal bottleneck in AI-driven drug discovery. Current scaffold-splitting protocols fail to obstruct microscopic semantic overlap, predisposing models to shortcut learning and overestimating their true extrapolation capability; meanwhile, conventional domain adaptation paradigms suffer under extreme structural shifts, as blindly aligning heterogeneous source libraries injects topological noise and triggers negative transfer. To address these two challenges, scaffold-cluster out-of-distribution performance evaluation benchmark (SCOPE-BENCH), a benchmark built on cluster-level partitioning in an explicit physicochemical descriptor space, is proposed alongside policy optimization for multi-source adaptation (POMA), a framework that formulates knowledge transfer as a retrieve-compose-adapt pipeline: labeled source scaffolds structurally close to the unlabeled target are first identified as proxy targets; a reinforcement learning policy then adaptively selects the optimal source subset from an exponentially large candidate pool; and dual-scale domain adaptation is finally performed at macroscopic topological and microscopic pharmacophore scales. Evaluations show that prediction errors of state-of-the-art 3D molecular models surge by up to 8.0x on SCOPE-BENCH with a mean of 5.9x, while POMA achieves up to an 11.2% reduction in mean absolute error with an average relative improvement of 6.2% across diverse backbone architectures. Code is available at https://anonymous.4open.science/r/Molecular-OOD-Code-73F6.
Original Article
View Cached Full Text

Cached at: 05/15/26, 06:25 AM

# Rethinking Molecular OOD Generalization via Target-Aware Source Selection
Source: [https://arxiv.org/html/2605.13932](https://arxiv.org/html/2605.13932)
\\@anonymousfalse

Zhuohao Lin†, Kun Li†, Jiameng Chen, Wenbin Hu∗ School of Computer Science Wuhan University Wuhan, China \{linzhuohao, likun98, jiameng\.chen, hwb\}@whu\.edu\.cn †These authors contributed equally\.∗Corresponding author\. Yizhen Zheng Department of Data Science and Artificial Intelligence, Monash University Victoria, Australia yizhen\.zheng1@monash\.edu &Jiajun Yu College of Computer Science and Technology, Zhejiang University Hangzhou, China jiajunyu1999@gmail\.com &Duanhua Cao School of Life Sciences and Technology, Tongji University Shanghai, 200092, China caodh@tongji\.edu\.cn

###### Abstract

Robust prediction of molecular properties under extreme out\-of\-distribution \(OOD\) scenarios is a pivotal bottleneck in AI\-driven drug discovery\. Current scaffold\-splitting protocols fail to obstruct microscopic semantic overlap, predisposing models to shortcut learning and overestimating their true extrapolation capability; meanwhile, conventional domain adaptation paradigms suffer under extreme structural shifts, as blindly aligning heterogeneous source libraries injects topological noise and triggers negative transfer\. To address these two challenges, scaffold\-cluster out\-of\-distribution performance evaluation benchmark \(SCOPE\-Bench\), a benchmark built on cluster\-level partitioning in an explicit physicochemical descriptor space, is proposed alongside policy optimization for multi\-source adaptation \(POMA\), a framework that formulates knowledge transfer as a retrieve–compose–adapt pipeline: labeled source scaffolds structurally close to the unlabeled target are first identified as proxy targets; a reinforcement learning policy then adaptively selects the optimal source subset from an exponentially large candidate pool; and dual\-scale domain adaptation is finally performed at macroscopic topological and microscopic pharmacophore scales\. Evaluations show that prediction errors of state\-of\-the\-art 3D molecular models surge by up to8\.0×8\.0\\timesonSCOPE\-Benchwith a mean of5\.9×5\.9\\times, whilePOMAachieves up to an11\.2%11\.2\\%reduction in mean absolute error with an average relative improvement of6\.2%6\.2\\%across diverse backbone architectures\. Code is available at[https://anonymous\.4open\.science/r/Molecular\-OOD\-Code\-73F6](https://anonymous.4open.science/r/Molecular-OOD-Code-73F6)\.

## 1Introduction

Modern drug discovery operates within a chemical space of approximately106010^\{60\}potential drug\-like molecules, where identifying lead compounds with specific biological activities\[[26](https://arxiv.org/html/2605.13932#bib.bib73)\]remains a fundamental challenge\[[12](https://arxiv.org/html/2605.13932#bib.bib9),[47](https://arxiv.org/html/2605.13932#bib.bib10)\]\. Molecular representation learning has become a cornerstone of this effort\[[38](https://arxiv.org/html/2605.13932#bib.bib11),[32](https://arxiv.org/html/2605.13932#bib.bib64),[57](https://arxiv.org/html/2605.13932#bib.bib13),[48](https://arxiv.org/html/2605.13932#bib.bib8)\], evolving from hand\-crafted descriptors\[[35](https://arxiv.org/html/2605.13932#bib.bib14)\]and 2D message passing networks\[[18](https://arxiv.org/html/2605.13932#bib.bib16),[14](https://arxiv.org/html/2605.13932#bib.bib17)\]to 3D geometric equivariant models such as ViSNet\[[49](https://arxiv.org/html/2605.13932#bib.bib19)\], ETNN\[[5](https://arxiv.org/html/2605.13932#bib.bib20)\], GotenNet\[[3](https://arxiv.org/html/2605.13932#bib.bib1)\], and SchNet\[[40](https://arxiv.org/html/2605.13932#bib.bib21)\], which capture high\-order geometric tensors while preserving physical consistency under rotation and translation\[[41](https://arxiv.org/html/2605.13932#bib.bib22),[37](https://arxiv.org/html/2605.13932#bib.bib23)\], approaching density functional theory accuracy on i\.i\.d\. benchmarks\[[16](https://arxiv.org/html/2605.13932#bib.bib26)\]\.

![Refer to caption](https://arxiv.org/html/2605.13932v1/x1.png)Figure 1:Core motivation of this work\. \(a\) Conventional scaffold splitting significantly overestimates model generalization due to underlying semantic overlap\. \(b\) Structural shifts under strict OOD settings trigger a multi\-fold surge in prediction errors\. \(c\) Source domain selection and composition act as the primary drivers of extrapolation performance, sometimes exceeding the impact of backbone architecture choice\.Despite this progress, a critical gap persists between benchmark performance and real\-world deployment\[[25](https://arxiv.org/html/2605.13932#bib.bib72),[27](https://arxiv.org/html/2605.13932#bib.bib27)\]\. In practice, novel molecules originate from structurally isolated regions of chemical space\[[20](https://arxiv.org/html/2605.13932#bib.bib29)\]\. Prevailing benchmarks rely on conventional scaffold splitting\[[7](https://arxiv.org/html/2605.13932#bib.bib32),[51](https://arxiv.org/html/2605.13932#bib.bib31)\], which mandates only that test scaffolds be absent from training\. However, this macroscopic decoupling fails to prevent microscopic semantic overlap\. Molecules with distinct scaffolds frequently share local conjugated pi\-electron systems or identical hydrogen\-bond donor networks\. This underlying overlap predisposes models to shortcut learning\[[13](https://arxiv.org/html/2605.13932#bib.bib33)\]rather than learning transferable physicochemical invariants\[[53](https://arxiv.org/html/2605.13932#bib.bib35)\]\. To address this fundamental evaluation bias, we propose the Scaffold\-cluster out\-of\-distribution performance evaluation benchmark \(SCOPE\-Bench\)\. This benchmark enforces strict metric separation based on explicit physicochemical descriptor clustering to completely preclude hidden structural interpolation\. As shown in Figure[1](https://arxiv.org/html/2605.13932#S1.F1)a and Figure[1](https://arxiv.org/html/2605.13932#S1.F1)b, models perform deceptively well under standard splits but collapse catastrophically under these stricter criteria\.

Extreme distribution shifts also expose the fatal limitations of existing transfer learning strategies\[[30](https://arxiv.org/html/2605.13932#bib.bib37),[31](https://arxiv.org/html/2605.13932#bib.bib38)\]\. Blindly aligning heterogeneous source libraries with an unlabeled target injects topological noise, which triggers dimensionality collapse\[[11](https://arxiv.org/html/2605.13932#bib.bib39)\]and severe negative transfer\[[36](https://arxiv.org/html/2605.13932#bib.bib40)\]\. Crucially, our preliminary analysis in Figure[1](https://arxiv.org/html/2605.13932#S1.F1)c reveals that the policy selection and composition of source domains act as the primary determinants of extrapolation success, often exerting a greater influence than the choice of backbone architecture itself\. This observation necessitates an intelligent mechanism to perceive the target domain and actively decide the optimal source configuration\. To tackle this challenge, we introduce policy optimization for multi\-source adaptation \(POMA\), which formulates knowledge transfer as an integrated, policy\-driven retrieve–compose–adapt pipeline\.

POMAintroduces two core innovations to distinguish its selection policy from conventional paradigms\. First, we revolutionize the retrieval and composition stages by learning a combinatorial selection policy via Group Relative Policy Optimization\[[42](https://arxiv.org/html/2605.13932#bib.bib42)\]\. Unlike static graph kernels that rank candidates in isolation, our policy dynamically explores an exponentially large combinatorial space to compose the most synergistic source subset without requiring a fragile value network\. Second, we innovate the domain adaptation stage by replacing conventional global alignment with a dual\-scale decoupled architecture\[[44](https://arxiv.org/html/2605.13932#bib.bib46)\]\. Standard adaptation methods force holistic feature matching, which inevitably destroys fine\-grained chemical semantics\.POMAresolves this by aligning macroscopic whole\-molecule topologies and microscopic pharmacophore fragments independently\. This dual\-scale regularization ensures that the selection policy is supported by a robust adaptation process that preserves both structural and chemical precision\.

Overall, our contributions can be summarized: \(i\) Scaffold\-cluster out\-of\-distribution performance evaluation benchmark \(SCOPE\-Bench\), a rigorous OOD benchmark based on physicochemical clustering that eliminates evaluation biases\. State\-of\-the\-art models degrade by up to8\.0×8\.0\\timeswith a mean of5\.9×5\.9\\times, exposing their fundamental OOD vulnerability\. \(ii\) Policy optimization for multi\-source adaptation \(POMA\), a policy\-guided framework that overcomes negative transfer through a target\-aware selection policy and dual\-scale decoupled domain adaptation\. \(iii\) Extensive experiments across 3D equivariant architectures demonstrate up to an11\.2%11\.2\\%reduction in mean absolute error with an average relative improvement of6\.2%6\.2\\%across all tasks, validating cross\-architecture universality\.

![Refer to caption](https://arxiv.org/html/2605.13932v1/x2.png)Figure 2:Overview of thePOMAframework as a retrieve–compose–adapt pipeline\. Labeled source scaffolds structurally close to the target are first identified as proxy targets to enable reward estimation\. A reinforcement learning policy then selects an optimal source subset from the candidate pool\. Finally, a dual\-scale adaptation module aligns macroscopic topologies and microscopic pharmacophore features, with policy updates driven by transfer performance on proxy tasks\.
## 2Related Work

Molecular machine learning\[[22](https://arxiv.org/html/2605.13932#bib.bib68),[21](https://arxiv.org/html/2605.13932#bib.bib79),[25](https://arxiv.org/html/2605.13932#bib.bib72)\]has become a core computational engine for drug discovery\[[9](https://arxiv.org/html/2605.13932#bib.bib78),[23](https://arxiv.org/html/2605.13932#bib.bib77)\], with its progress largely driven by increasingly faithful representations of molecular structure, from hand\-crafted fingerprints\[[35](https://arxiv.org/html/2605.13932#bib.bib14)\]to graph neural networks\[[55](https://arxiv.org/html/2605.13932#bib.bib66),[54](https://arxiv.org/html/2605.13932#bib.bib67)\]and, more recently, geometry\-aware equivariant models\[[52](https://arxiv.org/html/2605.13932#bib.bib71),[49](https://arxiv.org/html/2605.13932#bib.bib19),[3](https://arxiv.org/html/2605.13932#bib.bib1)\]\. Despite near\-density functional theory accuracy on i\.i\.d\. benchmarks\[[16](https://arxiv.org/html/2605.13932#bib.bib26)\], Hu et al\.\[[17](https://arxiv.org/html/2605.13932#bib.bib49)\]showed that this accuracy relies heavily on substructural overlap: once unseen topologies are encountered, generalization degrades severely\[[24](https://arxiv.org/html/2605.13932#bib.bib76)\]\.

Random splitting and MoleculeNet’s scaffold splitting\[[33](https://arxiv.org/html/2605.13932#bib.bib70),[51](https://arxiv.org/html/2605.13932#bib.bib31)\]progressively improved evaluation realism, yet neither prevents microscopic semantic overlap\. Models exploit spurious substructural correlations rather than causal physicochemical laws\[[13](https://arxiv.org/html/2605.13932#bib.bib33)\], causing severe calibration degradation under the dataset shifts\[[29](https://arxiv.org/html/2605.13932#bib.bib50),[38](https://arxiv.org/html/2605.13932#bib.bib11)\]\. The GOOD benchmark\[[15](https://arxiv.org/html/2605.13932#bib.bib52)\]and MoleOOD\[[53](https://arxiv.org/html/2605.13932#bib.bib35)\]demonstrate that without enforcing invariant risk minimization\[[1](https://arxiv.org/html/2605.13932#bib.bib53)\], shortcut learning leads researchers to systematically overestimate model robustness\.

Domain adaptation methods\[[8](https://arxiv.org/html/2605.13932#bib.bib54)\]based on MMD, adversarial training\[[56](https://arxiv.org/html/2605.13932#bib.bib58)\], or second\-order statistics\[[44](https://arxiv.org/html/2605.13932#bib.bib46)\]assume that merging all source domains and aligning them globally with a target is beneficial\. Under extreme molecular shifts, this assumption fails: heterogeneous source gradients overcompress the target representation, causing dimensionality collapse\[[11](https://arxiv.org/html/2605.13932#bib.bib39)\]and negative transfer\[[36](https://arxiv.org/html/2605.13932#bib.bib40),[50](https://arxiv.org/html/2605.13932#bib.bib41)\], making source selection a prerequisite for robust molecular adaptation\. Graph kernel methods\[[45](https://arxiv.org/html/2605.13932#bib.bib60)\]rank source candidates statically but cannot optimize based on downstream feedback\. Reinforcement learning\[[6](https://arxiv.org/html/2605.13932#bib.bib44)\]offers dynamic combinatorial selection, yet PPO\[[39](https://arxiv.org/html/2605.13932#bib.bib45)\]requires the same scale Critic network prone to collapse under sparse rewards\. GRPO\[[42](https://arxiv.org/html/2605.13932#bib.bib42)\]eliminates the value network entirely, computing advantages via intra\-group standardization, which naturally fits the problem of selecting the best source subset from a large candidate pool\.

## 3Method

### 3\.1Problem Definition

The prediction task is defined over 3D molecular graphsGi=\(𝒱i,ℰi,𝐑i,𝐙i\)G\_\{i\}=\(\\mathcal\{V\}\_\{i\},\\mathcal\{E\}\_\{i\},\\mathbf\{R\}\_\{i\},\\mathbf\{Z\}\_\{i\}\)with labelyi∈ℝy\_\{i\}\\in\\mathbb\{R\}; the topological scaffold of each molecule is extracted via the Bemis–Murcko functionsi=ϕ​\(Gi\)s\_\{i\}=\\phi\(G\_\{i\}\)\.

Unlike the i\.i\.d\. assumption of conventional empirical risk minimization, the extreme OOD setting requires that the source scaffold set𝒮\\mathcal\{S\}and unlabeled target set𝒯\\mathcal\{T\}be simultaneously disjoint \(𝒮∩𝒯=∅\\mathcal\{S\}\\cap\\mathcal\{T\}=\\emptyset\) and metrically separated: their 1\-Wasserstein distance in the physicochemical descriptor space must exceed a predefined thresholdτd​i​s​t\\tau\_\{dist\}, preventing shortcut learning via substructure interpolation\.

Two core challenges arise from this setting\.Challenge 1 \(Benchmark gap\)\.Even when scaffolds are disjoint, existing splits allowsupp⁡\(P​\(𝒮\)\)∩supp⁡\(P​\(𝒯\)\)≠∅\\operatorname\{supp\}\(P\(\\mathcal\{S\}\)\)\\cap\\operatorname\{supp\}\(P\(\\mathcal\{T\}\)\)\\neq\\emptyset, enabling models to exploit local feature overlap rather than achieving true extrapolation\.Challenge 2 \(Negative transfer\)\.Global alignment of the entire source pool𝒮\\mathcal\{S\}with target𝒯\\mathcal\{T\}injects heterogeneous gradients that over\-compress the target representation\. This causes dimensionality collapse, formally a drastic rank reduction of the target feature covariance matrix𝐂𝒯\\mathbf\{C\}\_\{\\mathcal\{T\}\}after alignment compared to the unaligned baseline, and leads to worse performance than using the optimal transferable subset𝒮∗⊂𝒮\\mathcal\{S\}^\{\*\}\\subset\\mathcal\{S\}alone\.

### 3\.2Construction ofSCOPE\-Bench

To achieve the metric separation required by the problem definition, scaffold\-cluster out\-of\-distribution performance evaluation benchmark \(SCOPE\-Bench\) constructs domains via a three\-step pipeline\.

Step 1: Scaffold extraction and feature construction\.Bemis–Murcko scaffolds are extracted from QM9\[[34](https://arxiv.org/html/2605.13932#bib.bib61)\]via RDKit\[[19](https://arxiv.org/html/2605.13932#bib.bib62)\], yielding 1,247 unique scaffolds with≥10\\geq 10samples each\. Each scaffoldsks\_\{k\}is represented by a four\-dimensional physicochemical feature vector:

fk=Vm​a​c​r​o⊕Ve​l​e​m​e​n​t⊕Vc​o​n​n⊕Vf​l​e​xf\_\{k\}=V\_\{macro\}\\oplus V\_\{element\}\\oplus V\_\{conn\}\\oplus V\_\{flex\}\(1\)where⊕\\oplusdenotes concatenation;Vm​a​c​r​oV\_\{macro\}encodes global topology \(atom/ring count\);Ve​l​e​m​e​n​tV\_\{element\}captures elemental polarizability;Vc​o​n​nV\_\{conn\}reflects electron delocalization and rigidity; andVf​l​e​xV\_\{flex\}measures conformational entropy via rotatable bond ratios\.

Step 2: Hierarchical clustering\.To handle the long\-tail distribution of polycyclic scaffolds, a hierarchical pre\-classification is applied: scaffolds are first grouped into five levels by ring count and maximum ring size, then feature vectors within each level are Z\-score normalized\. K\-Means\+\+\[[2](https://arxiv.org/html/2605.13932#bib.bib5)\]is applied within each level, with cluster quotas allocated proportionally to level size and adjusted by a sparsity compensation coefficientαc​o​m​p\\alpha\_\{comp\}for underrepresented polycyclic levels, yieldingKt​o​t​a​l=12K\_\{total\}=12globally disjoint clusters that form strict Voronoi boundaries in chemical space\.

Step 3: Asymmetric partitioning\.The clusters are partitioned asymmetrically to simulate a Universal Domain Adaptation task: a majority of clusters form the source domain, one cluster serves as the validation set, and the remaining clusters constitute the target domain\. A subset of target scaffolds with sufficient sample sizes is selected as independent zero\-shot extrapolation tasks, ensuring complete invisibility of target scaffolds during training\.

### 3\.3Policy Optimization for Multi\-source Adaptation

Under extreme structural shifts, forcibly aligning irrelevant heterogeneous scaffold distributions injects topological noise and induces dimensionality collapse\. To actively overcome negative transfer, we detail the policy optimization for multi\-source adaptation \(POMA\), which formulates molecular knowledge transfer as an integrated process of retrieve, compose, and adapt\. Unlike conventional static methods that rely on fixed similarity metrics, our framework learns an intelligent selection policy to actively explore the combinatorial space of source domains\. This policy\-centric design ensures that the model can perceive target structural features and decide the optimal knowledge transfer pathway to maximize extrapolation performance\.

#### 3\.3\.1Task\-Specific Environment Construction

In the unsupervised extreme UniDA task, directly applying reinforcement learning faces issues of reward sparsity and exponentially growing combinatorial search space\.To establish an effective gradient feedback path, Task\-Specific Environment Construction is proposed: scaffolds from the labeled source domains that are similar to the real target domain and possess broad generalization representativeness are selected to act as proxy targets\.

GivenNTN\_\{T\}unlabeled real targets𝒯r​e​a​l\\mathcal\{T\}\_\{real\}and a sufficiently large source domain pool𝒮p​o​o​l\\mathcal\{S\}\_\{pool\}, the Morgan fingerprint\[[28](https://arxiv.org/html/2605.13932#bib.bib2)\]cosine similarity between a candidate scaffolds∈𝒮p​o​o​ls\\in\\mathcal\{S\}\_\{pool\}and a real targetti∈𝒯r​e​a​lt\_\{i\}\\in\\mathcal\{T\}\_\{real\}is computed as:

Sim⁡\(s,ti\)=𝐯s⋅𝐯ti‖𝐯s‖⋅‖𝐯ti‖\+ϵ\\operatorname\{Sim\}\(s,t\_\{i\}\)=\\frac\{\\mathbf\{v\}\_\{s\}\\cdot\\mathbf\{v\}\_\{t\_\{i\}\}\}\{\\\|\\mathbf\{v\}\_\{s\}\\\|\\cdot\\\|\\mathbf\{v\}\_\{t\_\{i\}\}\\\|\+\\epsilon\}\(2\)where𝐯s\\mathbf\{v\}\_\{s\}and𝐯ti\\mathbf\{v\}\_\{t\_\{i\}\}denote the molecular fingerprint feature vectors of source scaffoldssand target scaffoldtit\_\{i\}, respectively,∥⋅∥\\\|\\cdot\\\|is theL2L\_\{2\}norm, andϵ\\epsilonis a small positive constant for numerical stability\. To comprehensively evaluate the suitability of a candidate scaffold as a proxy, a joint HubScore is defined:

HubScore⁡\(s\)=maxti∈𝒯r​e​a​l⁡Sim⁡\(s,ti\)\+λ​∑ti∈𝒯r​e​a​l𝕀​\(Sim⁡\(s,ti\)\>τs​i​m\)\\operatorname\{HubScore\}\(s\)=\\max\_\{t\_\{i\}\\in\\mathcal\{T\}\_\{real\}\}\\operatorname\{Sim\}\(s,t\_\{i\}\)\+\\lambda\\sum\_\{t\_\{i\}\\in\\mathcal\{T\}\_\{real\}\}\\mathbb\{I\}\\\!\\left\(\\operatorname\{Sim\}\(s,t\_\{i\}\)\>\\tau\_\{sim\}\\right\)\(3\)whereλ\>0\\lambda\>0is the coverage balancing coefficient,𝕀​\(⋅\)\\mathbb\{I\}\(\\cdot\)is the indicator function, andτs​i​m\\tau\_\{sim\}is the predefined similarity threshold\. The top\-NpN\_\{p\}scaffolds by HubScore become proxy targets𝒯p​r​o​x​y\\mathcal\{T\}\_\{proxy\}\. This ensures that an excellent proxy target maintains broad isomorphic connections with multiple real targets\. Critically, this screening process relies solely on unlabeled molecular fingerprint topological priors and does not involve any real property labels of the target domain\.

After determining the proxy targets, the Weisfeiler–Lehman \(WL\) graph kernel\[[43](https://arxiv.org/html/2605.13932#bib.bib63)\]kW​Lk\_\{WL\}combined with sample size information is used to rank candidate source domains:

Sr​a​n​k​\(s∗,cj\)=kW​L​\(s∗,cj\)×ln⁡\(1\+\|𝒟cj\|\)S\_\{rank\}\(s^\{\*\},c\_\{j\}\)=k\_\{WL\}\(s^\{\*\},c\_\{j\}\)\\times\\ln\(1\+\|\\mathcal\{D\}\_\{c\_\{j\}\}\|\)\(4\)wherekW​L​\(s∗,cj\)k\_\{WL\}\(s^\{\*\},c\_\{j\}\)is the WL kernel similarity between proxy targets∗s^\{\*\}and candidate domaincjc\_\{j\}, and\|𝒟cj\|\|\\mathcal\{D\}\_\{c\_\{j\}\}\|is the sample size ofcjc\_\{j\}\. The top\-MMcandidates by descendingSr​a​n​kS\_\{rank\}form the candidate pool𝒞\\mathcal\{C\}\.

#### 3\.3\.2Dual\-Scale Decoupled Domain Adaptation

We implement a dual\-scale decoupled alignment strategy to minimize the distribution discrepancy while preserving fine\-grained chemical semantics\. To ensure that the learned selection policy is supported by a robust adaptation process, the encoder constructs two parallel feature paths: a macroscopic path yielding whole\-molecule featureshm​o​l∈ℝdh\_\{mol\}\\in\\mathbb\{R\}^\{d\}and a microscopic path yielding pharmacophore featureshs​u​b∈ℝdh\_\{sub\}\\in\\mathbb\{R\}^\{d\}via BRICS retrosynthetic cleavage\[[10](https://arxiv.org/html/2605.13932#bib.bib59)\]\. For each path, the distribution gap between sourcekkand target is measured by the squared Frobenius norm of the difference between their empirical covariance matrices \(Deep CORAL\[[44](https://arxiv.org/html/2605.13932#bib.bib46)\]criterion\)\. Given theKKselected source domains with normalized transfer weightsγk=Sk/∑jSj\\gamma\_\{k\}=S\_\{k\}/\\sum\_\{j\}S\_\{j\}\(derived fromSr​a​n​kS\_\{rank\}\), the unified multi\-source alignment objective is:

ℒD​A=wm​o​l​∑k=1Kγk⋅‖𝐂sm​o​l,k−𝐂tm​o​l‖F24​d2\+ws​u​b​∑k=1Kγk⋅‖𝐂ss​u​b,k−𝐂ts​u​b‖F24​d2\\mathcal\{L\}\_\{DA\}=w\_\{mol\}\\sum\_\{k=1\}^\{K\}\\gamma\_\{k\}\\cdot\\frac\{\\\|\\mathbf\{C\}\_\{s\}^\{mol,k\}\-\\mathbf\{C\}\_\{t\}^\{mol\}\\\|\_\{F\}^\{2\}\}\{4d^\{2\}\}\+w\_\{sub\}\\sum\_\{k=1\}^\{K\}\\gamma\_\{k\}\\cdot\\frac\{\\\|\\mathbf\{C\}\_\{s\}^\{sub,k\}\-\\mathbf\{C\}\_\{t\}^\{sub\}\\\|\_\{F\}^\{2\}\}\{4d^\{2\}\}\(5\)where𝐂sm​o​l,k,𝐂tm​o​l∈ℝd×d\\mathbf\{C\}\_\{s\}^\{mol,k\},\\mathbf\{C\}\_\{t\}^\{mol\}\\in\\mathbb\{R\}^\{d\\times d\}are the empirical covariance matrices of the macroscopic features of source domainkkand target, respectively;𝐂ss​u​b,k\\mathbf\{C\}\_\{s\}^\{sub,k\}and𝐂ts​u​b\\mathbf\{C\}\_\{t\}^\{sub\}are the corresponding microscopic covariance matrices;∥⋅∥F\\\|\\cdot\\\|\_\{F\}denotes the Frobenius norm; andddis the feature dimension\. The total training objective combines supervised regression with alignment:

ℒt​o​t​a​l=wr​e​g​ℒr​e​g\+ℒD​A\\mathcal\{L\}\_\{total\}=w\_\{reg\}\\,\\mathcal\{L\}\_\{reg\}\+\\mathcal\{L\}\_\{DA\}\(6\)
Training proceeds in two phases\. For the firstEw​a​r​mE\_\{warm\}epochs, onlyℒr​e​g\\mathcal\{L\}\_\{reg\}is used for supervised warm\-up\. Subsequently, a dynamic weight controller adaptively balanceswr​e​gw\_\{reg\},wm​o​lw\_\{mol\}, andws​u​bw\_\{sub\}to ensure stable optimization\. Specifically,wr​e​gw\_\{reg\}is decayed as the source\-domain regression loss converges, whilewm​o​lw\_\{mol\}andws​u​bw\_\{sub\}are rebalanced via momentum\-smoothed updates with explicit clipping bounds\. This prevents either alignment scale from dominating under volatile training dynamics\.

#### 3\.3\.3Source domain combinatorial decision via GRPO

To encode candidate pool information into the decision network, a state representation vector𝐱j∈ℝds​t​a​t​e\\mathbf\{x\}\_\{j\}\\in\\mathbb\{R\}^\{d\_\{state\}\}withds​t​a​t​e=258d\_\{state\}=258is constructed for each candidate scaffoldcjc\_\{j\}:

𝐱j=fs∗⊕fcj⊕\[kW​L​\(s∗,cj\),ln⁡\(1\+\|𝒟cj\|\)cn​o​r​m\]\\mathbf\{x\}\_\{j\}=f\_\{s^\{\*\}\}\\oplus f\_\{c\_\{j\}\}\\oplus\\left\[k\_\{WL\}\(s^\{\*\},c\_\{j\}\),\\;\\frac\{\\ln\(1\+\|\\mathcal\{D\}\_\{c\_\{j\}\}\|\)\}\{c\_\{norm\}\}\\right\]\(7\)where⊕\\oplusdenotes vector concatenation,fs∗f\_\{s^\{\*\}\}andfcjf\_\{c\_\{j\}\}are the fingerprint features of the proxy target and candidate source, respectively, andcn​o​r​mc\_\{norm\}is a normalization constant for the logarithmic term\. The policy networkπθ\\pi\_\{\\theta\}employs a Multi\-Layer Perceptron with layer normalization\. For each candidate scaffold state𝐱j\\mathbf\{x\}\_\{j\}, the network outputs a confidence logit mapped via Sigmoid to a Bernoulli selection probabilitypjp\_\{j\}\. An action vector𝐚∈\{0,1\}M\\mathbf\{a\}\\in\\\{0,1\\\}^\{M\}is generated through independent Bernoulli sampling, with log\-likelihood:

ln⁡πθ​\(𝐚∣S\)=∑j=1M\[aj​ln⁡pj\+\(1−aj\)​ln⁡\(1−pj\)\]\\ln\\pi\_\{\\theta\}\(\\mathbf\{a\}\\mid S\)=\\sum\_\{j=1\}^\{M\}\\bigl\[a\_\{j\}\\ln p\_\{j\}\+\(1\-a\_\{j\}\)\\ln\(1\-p\_\{j\}\)\\bigr\]\(8\)
![Refer to caption](https://arxiv.org/html/2605.13932v1/fig3_1a.png)\(a\)Random split
![Refer to caption](https://arxiv.org/html/2605.13932v1/fig3_1b.png)\(b\)Standard scaffold split
![Refer to caption](https://arxiv.org/html/2605.13932v1/fig3_1c.png)\(c\)SCOPE\-Bench

Figure 3:t\-SNE feature distributions under different splitting protocols via Local Domain Dominance statistics\. Blue/red regions denote source/target dominant areas; white regions indicate feature overlap\. \(a\) Random split: domains are fully mixed\. \(b\) Standard scaffold split: partial clusters emerge, but significant overlap persists\. \(c\)SCOPE\-Bench: clear separation with a distributional vacuum zone, precluding interpolation shortcuts\.![Refer to caption](https://arxiv.org/html/2605.13932v1/fig3_2a.png)\(a\)Random split
![Refer to caption](https://arxiv.org/html/2605.13932v1/fig3_2b.png)\(b\)Standard scaffold split
![Refer to caption](https://arxiv.org/html/2605.13932v1/fig3_2c.png)\(c\)SCOPE\-Bench

Figure 4:Inter\-domain scaffold tanimoto similarity heatmaps\. Darker blue indicates higher similarity; lighter yellow indicates lower similarity\. \(a\) Random split: uniformly high similarity\. \(b\) Standard scaffold split: cross\-bands of high similarity persist, indicating unsevered substructure overlap\. \(c\)SCOPE\-Bench: similarity is strictly suppressed, achieving true substructure isolation\.For each GRPO iteration,GGaction sets are sampled in parallel; the reward for each isRi=MAEb​a​s​e−MAEiR\_\{i\}=\\mathrm\{MAE\}\_\{base\}\-\\mathrm\{MAE\}\_\{i\}measured after dual\-scale adaptation on𝒯p​r​o​x​y\\mathcal\{T\}\_\{proxy\}\. The intra\-group normalized advantage is:

A^i=Ri−μRVσRV\+ϵs\\hat\{A\}\_\{i\}=\\frac\{R\_\{i\}\-\\mu\_\{R\_\{V\}\}\}\{\\sigma\_\{R\_\{V\}\}\+\\epsilon\_\{s\}\}\(9\)whereμRV\\mu\_\{R\_\{V\}\}andσRV\\sigma\_\{R\_\{V\}\}are the mean and standard deviation of all sample rewards within the current group, andϵs\\epsilon\_\{s\}is a small constant for numerical stability\.

GRPO eliminates the need for a Critic network\. To constrain policy update step sizes, a reference networkπr​e​f\\pi\_\{ref\}is introduced, and the inverse KL divergence is estimated as a penalty\. LettingΔi=ln⁡πθ​\(𝐚i∣S\)−ln⁡πr​e​f​\(𝐚i∣S\)\\Delta\_\{i\}=\\ln\\pi\_\{\\theta\}\(\\mathbf\{a\}\_\{i\}\\mid S\)\-\\ln\\pi\_\{ref\}\(\\mathbf\{a\}\_\{i\}\\mid S\), the reverse KL divergence is approximated as:

DK​L​\(πθ∥πr​e​f\)≈1\|V\|​∑i∈V\(exp⁡\(Δi\)−Δi−1\)D\_\{KL\}\\\!\\left\(\\pi\_\{\\theta\}\\,\\\|\\,\\pi\_\{ref\}\\right\)\\approx\\frac\{1\}\{\|V\|\}\\sum\_\{i\\in V\}\\left\(\\exp\(\\Delta\_\{i\}\)\-\\Delta\_\{i\}\-1\\right\)\(10\)The GRPO optimization objective is:

ℒG​R​P​O​\(θ\)=−1\|V\|​∑i∈Vmin⁡\(ri​\(θ\)​A^i,clip⁡\(ri​\(θ\),1−ϵc​l​i​p,1\+ϵc​l​i​p\)​A^i\)\+β​DK​L​\(πθ∥πr​e​f\)\\mathcal\{L\}\_\{GRPO\}\(\\theta\)=\-\\frac\{1\}\{\|V\|\}\\sum\_\{i\\in V\}\\min\\\!\\left\(r\_\{i\}\(\\theta\)\\,\\hat\{A\}\_\{i\},\\;\\operatorname\{clip\}\\\!\\left\(r\_\{i\}\(\\theta\),\\,1\-\\epsilon\_\{clip\},\\,1\+\\epsilon\_\{clip\}\\right\)\\hat\{A\}\_\{i\}\\right\)\+\\beta\\,D\_\{KL\}\\\!\\left\(\\pi\_\{\\theta\}\\,\\\|\\,\\pi\_\{ref\}\\right\)\(11\)whereri=πθ​\(𝐚i∣S\)/πr​e​f​\(𝐚i∣S\)r\_\{i\}=\\pi\_\{\\theta\}\(\\mathbf\{a\}\_\{i\}\\mid S\)/\\pi\_\{ref\}\(\\mathbf\{a\}\_\{i\}\\mid S\)is the probability ratio,ϵc​l​i​p\\epsilon\_\{clip\}is the clipping threshold, andβ\>0\\beta\>0is the KL penalty coefficient\. At inference,πθ\\pi\_\{\\theta\}selects𝒮∗\\mathcal\{S\}^\{\*\}via a single forward pass;fθf\_\{\\theta\}is then fine\-tuned on𝒮∗\\mathcal\{S\}^\{\*\}withℒt​o​t​a​l\\mathcal\{L\}\_\{total\}for zero\-shot inference on𝒯r​e​a​l\\mathcal\{T\}\_\{real\}\. The full training procedure is detailed in Algorithm[1](https://arxiv.org/html/2605.13932#alg1)in the appendix\.

Table 1:Overall mean absolute error measured in eV on the QM9 dataset under different split protocols\. The table illustrates the catastrophic degradation from the standard scaffold split to theSCOPE\-Benchsupervised baseline and the subsequent performance recovery achieved byPOMA\. The third row quantifies the performance drop as a degradation factor\. The last row in red indicates the relative improvement achieved by our framework\.

### 3\.4Complexity Analysis

LetNNdenote the total source molecules,VVthe average atoms per molecule,MMthe candidate pool size, andNKN\_\{\\text\{K\}\}the average samples per selected domain, while treating other hyperparameters as fixed constants\. The framework decouples into three phases\. Offline preprocessing requires a one\-time computational cost of𝒪​\(N⋅V\)\\mathcal\{O\}\(N\\cdot V\)for fingerprint and index construction, with on\-disk storage scaling as𝒪​\(N\)\\mathcal\{O\}\(N\)\. During GRPO policy optimization, fixed\-length proxy adaptations and sparse message passing ensure each action evaluation incurs𝒪​\(V\)\\mathcal\{O\}\(V\)time, while WL kernel ranking adds𝒪​\(M\)\\mathcal\{O\}\(M\), resulting in a total GRPO time of𝒪​\(TGRPO⋅\(V\+M\)\)\\mathcal\{O\}\(T\_\{\\text\{GRPO\}\}\\cdot\(V\+M\)\)\. Once the candidate pool has been constructed and fixed, the GRPO optimization cost no longer scales with the total number of source moleculesNN\.

## 4Experiments

### 4\.1Experimental Setup

SCOPE\-Benchis built on QM9\[[34](https://arxiv.org/html/2605.13932#bib.bib61)\]following the protocol of Section[3\.2](https://arxiv.org/html/2605.13932#S3.SS2)\. Three quantum chemical properties are evaluated: the highest occupied molecular orbital energy \(HOMO\), the lowest unoccupied molecular orbital energy \(LUMO\), and HOMO–LUMO gap \(Gap\), all in eV\. All results are obtained with a fixed random seed of 42 for reproducibility\. Two protocols are compared: a supervised baseline fine\-tuned on the source domain for 200 epochs without adaptation, andPOMAwith offline policy optimization followed by target\-specific adaptation\.

### 4\.2Verification of SCOPE\-Bench

SCOPE\-Benchpartitions 133,885 QM9 molecules into 12 disjoint scaffold clusters: 6 form the source domain \(94,562 samples\), 1 the validation set \(18,326 samples\), and 5 the target domain \(19,894 samples\), with 15 target scaffolds ofN≥200N\\geq 200serving as independent zero\-shot tasks\. T\-SNE\[[46](https://arxiv.org/html/2605.13932#bib.bib6)\]using local domain dominance statistics with a threshold of eighty percent reveals that random and scaffold splits produce heavily overlapping latent distributions\. Conversely,SCOPE\-Benchenforces sharply separated and non\-overlapping manifolds with a clear distributional vacuum zone as shown in Figure[3](https://arxiv.org/html/2605.13932#S3.F3)\. Pairwise tanimoto similarity\[[4](https://arxiv.org/html/2605.13932#bib.bib7)\]computed via Morgan fingerprints shows that conventional splits retain dense high\-similarity cross\-domain patches, whereasSCOPE\-Benchsuppresses pairs with tanimoto above 0\.5 to near zero, confirming true microscopic orthogonality as shown in Figure[4](https://arxiv.org/html/2605.13932#S3.F4)\. As shown in the top section of Table[1](https://arxiv.org/html/2605.13932#S3.T1), all three backbones suffer catastrophic MAE increases onSCOPE\-Bench, with a mean degradation of5\.9×5\.9\\times, demonstrating that current state\-of\-the\-art architectures lack genuine OOD extrapolation capability\.

Table 2:Ablation study on ViSNet for the HOMO task usingSCOPE\-Bench\. The upper block compares source selection policies with full dual\-scale adaptation enabled\. The lower block shows alignment module ablation with the GRPO\-learned selection policy fixed\. A positive value ofΔ\\Deltaindicates improvement over the supervised baseline, while a negative value represents a decrease in performance\. The checkmark \(✓\\checkmark\) indicates that the corresponding module is enabled, while the cross \(×\\times\) indicates it is disabled\.PolicyMol\-CORAL \(M\)Sub\-CORAL \(S\)MAEΔ\\DeltaOnly ViSNet\-\-0\.1621\-Random selection✓\\checkmark✓\\checkmark0\.1708−5\.3%\-5\.3\\%Shallow feature matching✓\\checkmark✓\\checkmark0\.1663−2\.5%\-2\.5\\%Deep feature matching✓\\checkmark✓\\checkmark0\.1669−2\.9%\-2\.9\\%Physical descriptor✓\\checkmark✓\\checkmark0\.1641−1\.2%\-1\.2\\%Graph kernel similarity✓\\checkmark✓\\checkmark0\.1616\+0\.3%\+0\.3\\%Mixed strategy✓\\checkmark✓\\checkmark0\.1576\+2\.8%\+2\.8\\%POMAw/o M×\\times✓\\checkmark0\.1689−4\.2%\-4\.2\\%POMAw/o S✓\\checkmark×\\times0\.1672−3\.1%\-3\.1\\%POMA\(Full\)✓\\checkmark✓\\checkmark0\.1540\+5\.0%\\mathbf\{\+5\.0\\%\}![Refer to caption](https://arxiv.org/html/2605.13932v1/x3.png)\(a\)MAE vs\. subset sizeKKatM=50M=50
![Refer to caption](https://arxiv.org/html/2605.13932v1/x4.png)\(b\)MAE vs\. pool sizeMMatK=5K=5

Figure 5:Hyperparameter sensitivity on the HOMO task for ViSNet usingSCOPE\-Bench\. Shaded regions show variation across the 15 target scaffolds\.
### 4\.3Overall Results

To assess whetherPOMAeffectively overcomes negative transfer under extreme structural shifts, a zero\-shot extrapolation evaluation is conducted on the 15 target scaffolds ofSCOPE\-Bench, built on QM9, across the three backbone architectures\. The goal is to verify that target\-aware source selection consistently reduces mean absolute error compared to the supervised baseline that uses all available source data without adaptation\.Table[1](https://arxiv.org/html/2605.13932#S3.T1)reports the mean absolute error aggregated over fifteen tasks while the complete per\-scaffold performance breakdown for each property\-architecture combination is provided in Appendix[3](https://arxiv.org/html/2605.13932#A2.T3)\. This granular presentation ensures a comprehensive evaluation and excludes the possibility of coincidental performance on specific structural clusters\.

As shown in the bottom section of Table[1](https://arxiv.org/html/2605.13932#S3.T1),POMAachieves consistent improvements across all properties\. For HOMO, mean absolute error is reduced by 5\.0%, 8\.6%, and 2\.8% on ViSNet, ETNN, and GotenNet, respectively, demonstrating strong cross\-architecture transferability\. For LUMO, reductions of 8\.5%, 11\.2%, and 9\.1% are observed on ViSNet, ETNN, and GotenNet\. For the HOMO\-LUMO Gap, all three models achieve positive error reductions of 1\.4%, 7\.3%, and 1\.8%, respectively\. This confirms that target\-aware source selection successfully mitigates the sensitivity typically introduced by dual\-energy subtraction under extreme distribution shifts\. The empirical results validate that the proposed selection policy successfully mitigates the catastrophic degradation observed onSCOPE\-Bench\. By identifying synergistic source combinations as suggested in Figure[1](https://arxiv.org/html/2605.13932#S1.F1)c, the learned policy achieves a robust performance recovery across all evaluated architectures\. It inherits backbone hyperparameters without backbone\-specific tuning to demonstrate a powerful plug\-and\-play nature\.

### 4\.4Ablation Study

The GRPO\-based source retrieval and dual\-scale adaptation are the two core components ofPOMA\.

Q1: Does the RL\-based retrieval outperform heuristic source selection?To answer Q1, six non\-RL heuristic variants are compared on ViSNet with the full adaptation module fixed\. Random selection simulates blind alignment; Shallow\-feat\. uses cosine similarity of 40\-epoch graph embeddings; Deep\-feat\. uses 200\-epoch embeddings; Physical ranks by atomic composition and orbital energy penalties; Graph\-kernel uses WL kernel similarity; and Mixed\-strategy combines WL kernel and physical scores linearly\. As shown in the upper block of Table[2](https://arxiv.org/html/2605.13932#S4.T2), Deep\-feat\. triggers negative transfer, confirming that fully converged source embeddings suffer manifold distortion\.POMAsignificantly outperforms the strongest heuristic \(\+5\.0%\+5\.0\\%vs\.\+2\.8%\+2\.8\\%\), confirmed by a pairedtt\-test across all 15 extrapolation tasks \(p<0\.05p<0\.05\)\.

Q2: How do the alignment modules contribute independently?As shown in the lower block of Table[2](https://arxiv.org/html/2605.13932#S4.T2), removing either alignment module causes performance to fall*below*the supervised baseline \(−4\.2%\-4\.2\\%and−3\.1%\-3\.1\\%, respectively\), revealing that source selection alone without distribution alignment can introduce noise that actively harms generalization\. This confirms that both Mol\-CORAL and Sub\-CORAL are essential: macroscopic topology alignment reduces covariate shift at the scaffold level, while Sub\-CORAL at the pharmacophore level captures complementary local structural semantics that are invisible to global covariance alignment\. The fullPOMAmodel, integrating both components, achieves\+5\.0%\+5\.0\\%, validating their synergistic contribution\. We verify the contribution of each component across all fifteen target scaffolds to ensure that the observed performance gains are not specific to certain structural clusters\. Appendix[4](https://arxiv.org/html/2605.13932#A3.T4)provides the exhaustive per\-scaffold ablation results for all tasks, which confirms the robustness of the synergistic effect between the selection policy and decoupled alignment\.

### 4\.5Hyperparameter Analysis

The sensitivity ofPOMAto candidate pool sizeMMand subset sizeKKis evaluated on the HOMO task using the ViSNet architecture\. Figure[5\(a\)](https://arxiv.org/html/2605.13932#S4.F5.sf1)reveals a threshold effect for the subset sizeKK\. Specifically, withM=50M=50fixed, MAE remains stagnant fromK=1K=1toK=3K=3but drops significantly atK=5K=5\. This indicates that a critical mass of complementary source scaffolds is necessary to bridge the extreme distribution gap, because insufficient sources fail to provide adequate transferable invariants\. Figure[5\(b\)](https://arxiv.org/html/2605.13932#S4.F5.sf2)shows that withK=5K=5fixed, too small anMMrestricts policy exploration, while too large anMMdilutes the reward signal with low\-quality candidates\. Consequently,M=50M=50achieves the optimal balance and is adopted throughout\.

Furthermore, the default setting of 40 GRPO steps with group sizeG=33G=33already achieves optimal performance\. Scaling to3×3\\timestraining steps yields only marginal improvement at significantly higher computational cost\. This confirms that the intra\-group relative advantage mechanism provides efficient policy gradients even in small\-budget regimes\.

## 5Conclusion

In this paper, we introducedSCOPE\-Benchto resolve the pervasive issue of microscopic semantic overlap in molecular benchmarks\. Furthermore, we developedPOMAas a policy\-driven pipeline that revolutionizes knowledge transfer through combinatorial source selection and dual\-scale decoupled domain adaptation\. The learned policy dynamically evaluates the transfer value of source scaffolds via group relative policy optimization to effectively overcome negative transfer and dimensionality collapse caused by blind global alignment\. Extensive experiments onSCOPE\-Benchdemonstrate that prediction errors of state\-of\-the\-art models surge by a mean of5\.9×5\.9\\timescompared to conventional scaffold splits, whilePOMAachieves an average performance gain of6\.2%6\.2\\%over the supervised fine\-tuning baseline\. A current limitation is the computational cost of the GRPO training loop, which requires executing multiple proxy domain adaptation procedures per policy iteration\. In future work, surrogate reward models will be explored to amortize this cost, and the framework will be extended to zero\-proxy target scenarios\.

## References

- \[1\]\(2019\)Invariant risk minimization\.arXiv preprint arXiv:1907\.02893\.Cited by:[§2](https://arxiv.org/html/2605.13932#S2.p2.1)\.
- \[2\]D\. Arthur, S\. Vassilvitskii,et al\.\(2007\)K\-means\+\+: the advantages of careful seeding\.InSoda,Vol\.7,pp\. 1027–1035\.Cited by:[§3\.2](https://arxiv.org/html/2605.13932#S3.SS2.p3.2)\.
- \[3\]S\. Aykent and T\. Xia\(2025\)Gotennet: rethinking efficient 3d equivariant graph neural networks\.InThe Thirteenth International Conference on Learning Representations,Cited by:[Table 3](https://arxiv.org/html/2605.13932#A2.T3.9.9.10.1.5),[§1](https://arxiv.org/html/2605.13932#S1.p1.1),[§2](https://arxiv.org/html/2605.13932#S2.p1.1),[Table 1](https://arxiv.org/html/2605.13932#S3.T1.9.9.10.1.4)\.
- \[4\]D\. Bajusz, A\. Rácz, and K\. Héberger\(2015\)Why is tanimoto index an appropriate choice for fingerprint\-based similarity calculations?\.Journal of cheminformatics7\(1\),pp\. 20\.Cited by:[§4\.2](https://arxiv.org/html/2605.13932#S4.SS2.p1.2)\.
- \[5\]C\. Battiloro, M\. Tec, G\. Dasoulas, M\. Audirac, F\. Dominici,et al\.\(2024\)E \(n\) equivariant topological neural networks\.arXiv preprint arXiv:2405\.15429\.Cited by:[Table 3](https://arxiv.org/html/2605.13932#A2.T3.9.9.10.1.4),[§1](https://arxiv.org/html/2605.13932#S1.p1.1),[Table 1](https://arxiv.org/html/2605.13932#S3.T1.9.9.10.1.3)\.
- \[6\]I\. Bello, H\. Pham, Q\. V\. Le, M\. Norouzi, and S\. Bengio\(2016\)Neural combinatorial optimization with reinforcement learning\.arXiv preprint arXiv:1611\.09940\.Cited by:[§2](https://arxiv.org/html/2605.13932#S2.p3.1)\.
- \[7\]G\. W\. Bemis and M\. A\. Murcko\(1996\)The properties of known drugs\. 1\. molecular frameworks\.Journal of medicinal chemistry39\(15\),pp\. 2887–2893\.Cited by:[§1](https://arxiv.org/html/2605.13932#S1.p2.1)\.
- \[8\]S\. Ben\-David, J\. Blitzer, K\. Crammer, A\. Kulesza, F\. Pereira, and J\. W\. Vaughan\(2010\)A theory of learning from different domains\.Machine learning79\(1\),pp\. 151–175\.Cited by:[§2](https://arxiv.org/html/2605.13932#S2.p3.1)\.
- \[9\]X\. Chen, R\. Wu, Y\. Lan, T\. Ma, and Y\. Liu\(2026\)MolEvolve: llm\-guided evolutionary search for interpretable molecular optimization\.arXiv preprint arXiv:2603\.24382\.Cited by:[§2](https://arxiv.org/html/2605.13932#S2.p1.1)\.
- \[10\]J\. Degen, C\. Wegscheid\-Gerlach, A\. Zaliani, and M\. Rarey\(2008\)On the art of compiling and using’drug\-like’chemical fragment spaces\.ChemMedChem3\(10\),pp\. 1503\.Cited by:[§3\.3\.2](https://arxiv.org/html/2605.13932#S3.SS3.SSS2.p1.6)\.
- \[11\]H\. Fang, P\. Lu, and H\. Lin\(2024\)Tackling dimensional collapse toward comprehensive universal domain adaptation\.arXiv preprint arXiv:2410\.11271\.Cited by:[§1](https://arxiv.org/html/2605.13932#S1.p3.1),[§2](https://arxiv.org/html/2605.13932#S2.p3.1)\.
- \[12\]A\. Gaulton, L\. J\. Bellis, A\. P\. Bento, J\. Chambers, M\. Davies, A\. Hersey, Y\. Light, S\. McGlinchey, D\. Michalovich, B\. Al\-Lazikani,et al\.\(2012\)ChEMBL: a large\-scale bioactivity database for drug discovery\.Nucleic acids research40\(D1\),pp\. D1100–D1107\.Cited by:[§1](https://arxiv.org/html/2605.13932#S1.p1.1)\.
- \[13\]R\. Geirhos, J\. Jacobsen, C\. Michaelis, R\. Zemel, W\. Brendel, M\. Bethge, and F\. A\. Wichmann\(2020\)Shortcut learning in deep neural networks\.Nature Machine Intelligence2\(11\),pp\. 665–673\.Cited by:[§1](https://arxiv.org/html/2605.13932#S1.p2.1),[§2](https://arxiv.org/html/2605.13932#S2.p2.1)\.
- \[14\]J\. Gilmer, S\. S\. Schoenholz, P\. F\. Riley, O\. Vinyals, and G\. E\. Dahl\(2017\)Neural message passing for quantum chemistry\.InInternational conference on machine learning,pp\. 1263–1272\.Cited by:[§1](https://arxiv.org/html/2605.13932#S1.p1.1)\.
- \[15\]S\. Gui, X\. Li, L\. Wang, and S\. Ji\(2022\)Good: a graph out\-of\-distribution benchmark\.Advances in Neural Information Processing Systems35,pp\. 2059–2073\.Cited by:[§2](https://arxiv.org/html/2605.13932#S2.p2.1)\.
- \[16\]P\. Hohenberg and W\. Kohn\(1964\)Inhomogeneous electron gas\.Physical review136\(3B\),pp\. B864\.Cited by:[§1](https://arxiv.org/html/2605.13932#S1.p1.1),[§2](https://arxiv.org/html/2605.13932#S2.p1.1)\.
- \[17\]W\. Hu, B\. Liu, J\. Gomes, M\. Zitnik, P\. Liang, V\. Pande, and J\. Leskovec\(2019\)Strategies for pre\-training graph neural networks\.arXiv preprint arXiv:1905\.12265\.Cited by:[§2](https://arxiv.org/html/2605.13932#S2.p1.1)\.
- \[18\]T\. N\. Kipf and M\. Welling\(2016\)Semi\-supervised classification with graph convolutional networks\.arXiv preprint arXiv:1609\.02907\.Cited by:[§1](https://arxiv.org/html/2605.13932#S1.p1.1)\.
- \[19\]G\. Landrumet al\.\(2024\)RDKit: open\-source chemoinformatics\.Zenodo\.External Links:[Document](https://dx.doi.org/10.5281/zenodo.12782092),[Link](https://doi.org/10.5281/zenodo.12782092)Cited by:[§3\.2](https://arxiv.org/html/2605.13932#S3.SS2.p2.2)\.
- \[20\]H\. Li, X\. Wang, Z\. Zhang, and W\. Zhu\(2025\)Out\-of\-distribution generalization on graphs: a survey\.IEEE Transactions on Pattern Analysis and Machine Intelligence\.Cited by:[§1](https://arxiv.org/html/2605.13932#S1.p2.1)\.
- \[21\]K\. Li, L\. Hu, J\. Chen, H\. Zhang, Y\. Xiong, X\. Cai, W\. Hu, and J\. Wu\(2026\)Can molecular evolution mechanism enhance molecular representation?\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 15108–15116\.Cited by:[§2](https://arxiv.org/html/2605.13932#S2.p1.1)\.
- \[22\]K\. Li, L\. Hu, Y\. Xiong, J\. Yu, H\. Zhang, J\. Chen, X\. Cai, J\. Wu, and W\. Hu\(2026\)PCEvo: path\-consistent molecular representation via virtual evolutionary\.InProceedings of the Thirty\-Fourth International Joint Conference on Artificial Intelligence, IJCAI\-26,Note:Main TrackCited by:[§2](https://arxiv.org/html/2605.13932#S2.p1.1)\.
- \[23\]K\. Li, Z\. Wu, S\. Wang, J\. Wu, S\. Pan, and W\. Hu\(2025\)DrugPilot: llm\-based parameterized reasoning agent for drug discovery\.arXiv preprint arXiv:2505\.13940\.Cited by:[§2](https://arxiv.org/html/2605.13932#S2.p1.1)\.
- \[24\]K\. Li, Z\. Wu, Y\. Xiong, H\. Zhang, L\. Hu, Z\. Liu, J\. Zeng, W\. Wu, M\. Chen, J\. Chen,et al\.\(2025\)BSL: a unified and generalizable multitask learning platform for virtual drug discovery from design to synthesis\.arXiv preprint arXiv:2508\.01195\.Cited by:[§2](https://arxiv.org/html/2605.13932#S2.p1.1)\.
- \[25\]K\. Li, Y\. Xiong, H\. Zhang, X\. Cai, J\. Wu, B\. Du, and W\. Hu\(2025\)Graph\-structured small molecule drug discovery through deep learning: progress, challenges, and opportunities\.In2025 IEEE International Conference on Web Services \(ICWS\),Vol\.,pp\. 1033–1042\.External Links:[Document](https://dx.doi.org/10.1109/ICWS67624.2025.00135)Cited by:[§1](https://arxiv.org/html/2605.13932#S1.p2.1),[§2](https://arxiv.org/html/2605.13932#S2.p1.1)\.
- \[26\]K\. Li, Y\. Zeng, Y\. Xiong, H\. Wu, S\. Fang, Z\. Qu, Y\. Zhu, B\. Du, Z\. Gao, and W\. Hu\(2025\)Contrastive learning\-based drug screening model for glun1/glun3a inhibitors\.Acta Pharmacologica Sinica,pp\. 1–13\.Cited by:[§1](https://arxiv.org/html/2605.13932#S1.p1.1)\.
- \[27\]J\. Liu, Z\. Shen, Y\. He, X\. Zhang, R\. Xu, H\. Yu, and P\. Cui\(2021\)Towards out\-of\-distribution generalization: a survey\.arXiv preprint arXiv:2108\.13624\.Cited by:[§1](https://arxiv.org/html/2605.13932#S1.p2.1)\.
- \[28\]H\. L\. Morgan\(1965\)The generation of a unique machine description for chemical structures\-a technique developed at chemical abstracts service\.\.Journal of chemical documentation5\(2\),pp\. 107–113\.Cited by:[§3\.3\.1](https://arxiv.org/html/2605.13932#S3.SS3.SSS1.p2.5)\.
- \[29\]Y\. Ovadia, E\. Fertig, J\. Ren, Z\. Nado, D\. Sculley, S\. Nowozin, J\. Dillon, B\. Lakshminarayanan, and J\. Snoek\(2019\)Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift\.Advances in neural information processing systems32\.Cited by:[§2](https://arxiv.org/html/2605.13932#S2.p2.1)\.
- \[30\]S\. J\. Pan and Q\. Yang\(2009\)A survey on transfer learning\.IEEE Transactions on knowledge and data engineering22\(10\),pp\. 1345–1359\.Cited by:[§1](https://arxiv.org/html/2605.13932#S1.p3.1)\.
- \[31\]X\. Peng, Q\. Bai, X\. Xia, Z\. Huang, K\. Saenko, and B\. Wang\(2019\)Moment matching for multi\-source domain adaptation\.InProceedings of the IEEE/CVF international conference on computer vision,pp\. 1406–1415\.Cited by:[§1](https://arxiv.org/html/2605.13932#S1.p3.1)\.
- \[32\]M\. Popova, O\. Isayev, and A\. Tropsha\(2018\)Deep reinforcement learning for de novo drug design\.Science advances4\(7\),pp\. eaap7885\.Cited by:[§1](https://arxiv.org/html/2605.13932#S1.p1.1)\.
- \[33\]X\. Qin, C\. Wang, Z\. Zhou, L\. Chen, W\. Du, and Y\. Wang\(2026\)MSAnchor: de novo molecular generation from mass spectrometry data with anchor\-extended molecular scaffolds\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 953–961\.Cited by:[§2](https://arxiv.org/html/2605.13932#S2.p2.1)\.
- \[34\]R\. Ramakrishnan, P\. O\. Dral, M\. Rupp, and O\. A\. Von Lilienfeld\(2014\)Quantum chemistry structures and properties of 134 kilo molecules\.Scientific data1\(1\),pp\. 1–7\.Cited by:[§3\.2](https://arxiv.org/html/2605.13932#S3.SS2.p2.2),[§4\.1](https://arxiv.org/html/2605.13932#S4.SS1.p1.1)\.
- \[35\]D\. Rogers and M\. Hahn\(2010\)Extended\-connectivity fingerprints\.Journal of chemical information and modeling50\(5\),pp\. 742–754\.Cited by:[§1](https://arxiv.org/html/2605.13932#S1.p1.1),[§2](https://arxiv.org/html/2605.13932#S2.p1.1)\.
- \[36\]M\. T\. Rosenstein, Z\. Marx, L\. P\. Kaelbling, and T\. G\. Dietterich\(2005\)To transfer or not to transfer\.InNIPS Workshop on Transfer Learning,Cited by:[§1](https://arxiv.org/html/2605.13932#S1.p3.1),[§2](https://arxiv.org/html/2605.13932#S2.p3.1)\.
- \[37\]V\. G\. Satorras, E\. Hoogeboom, and M\. Welling\(2021\)E \(n\) equivariant graph neural networks\.InInternational conference on machine learning,pp\. 9323–9332\.Cited by:[§1](https://arxiv.org/html/2605.13932#S1.p1.1)\.
- \[38\]G\. Scalia, C\. A\. Grambow, B\. Pernici, Y\. Li, and W\. H\. Green\(2020\)Evaluating scalable uncertainty estimation methods for deep learning\-based molecular property prediction\.Journal of chemical information and modeling60\(6\),pp\. 2697–2717\.Cited by:[§1](https://arxiv.org/html/2605.13932#S1.p1.1),[§2](https://arxiv.org/html/2605.13932#S2.p2.1)\.
- \[39\]J\. Schulman, F\. Wolski, P\. Dhariwal, A\. Radford, and O\. Klimov\(2017\)Proximal policy optimization algorithms\.arXiv preprint arXiv:1707\.06347\.Cited by:[§2](https://arxiv.org/html/2605.13932#S2.p3.1)\.
- \[40\]K\. Schütt, P\. Kindermans, H\. E\. Sauceda Felix, S\. Chmiela, A\. Tkatchenko, and K\. Müller\(2017\)Schnet: a continuous\-filter convolutional neural network for modeling quantum interactions\.Advances in neural information processing systems30\.Cited by:[§1](https://arxiv.org/html/2605.13932#S1.p1.1)\.
- \[41\]K\. Schütt, O\. Unke, and M\. Gastegger\(2021\)Equivariant message passing for the prediction of tensorial properties and molecular spectra\.InInternational conference on machine learning,pp\. 9377–9388\.Cited by:[§1](https://arxiv.org/html/2605.13932#S1.p1.1)\.
- \[42\]Z\. Shao, P\. Wang, Q\. Zhu, R\. Xu, J\. Song, X\. Bi, H\. Zhang, M\. Zhang, Y\. Li, Y\. Wu,et al\.\(2024\)Deepseekmath: pushing the limits of mathematical reasoning in open language models\.arXiv preprint arXiv:2402\.03300\.Cited by:[§1](https://arxiv.org/html/2605.13932#S1.p4.1),[§2](https://arxiv.org/html/2605.13932#S2.p3.1)\.
- \[43\]N\. Shervashidze, P\. Schweitzer, E\. J\. Van Leeuwen, K\. Mehlhorn, and K\. M\. Borgwardt\(2011\)Weisfeiler\-lehman graph kernels\.\.Journal of Machine Learning Research12\(9\)\.Cited by:[§3\.3\.1](https://arxiv.org/html/2605.13932#S3.SS3.SSS1.p3.1)\.
- \[44\]B\. Sun and K\. Saenko\(2016\)Deep coral: correlation alignment for deep domain adaptation\.InEuropean conference on computer vision,pp\. 443–450\.Cited by:[§1](https://arxiv.org/html/2605.13932#S1.p4.1),[§2](https://arxiv.org/html/2605.13932#S2.p3.1),[§3\.3\.2](https://arxiv.org/html/2605.13932#S3.SS3.SSS2.p1.6)\.
- \[45\]M\. Togninalli, E\. Ghisu, F\. Llinares\-López, B\. Rieck, and K\. Borgwardt\(2019\)Wasserstein weisfeiler\-lehman graph kernels\.Advances in neural information processing systems32\.Cited by:[§2](https://arxiv.org/html/2605.13932#S2.p3.1)\.
- \[46\]L\. Van der Maaten and G\. Hinton\(2008\)Visualizing data using t\-sne\.\.Journal of machine learning research9\(11\)\.Cited by:[§4\.2](https://arxiv.org/html/2605.13932#S4.SS2.p1.2)\.
- \[47\]W\. P\. Walters and M\. Murcko\(2020\)Assessing the impact of generative ai on medicinal chemistry\.Nature biotechnology38\(2\),pp\. 143–145\.Cited by:[§1](https://arxiv.org/html/2605.13932#S1.p1.1)\.
- \[48\]H\. Wang, T\. Fu, Y\. Du, W\. Gao, K\. Huang, Z\. Liu, P\. Chandak, S\. Liu, P\. Van Katwyk, A\. Deac,et al\.\(2023\)Scientific discovery in the age of artificial intelligence\.Nature620\(7972\),pp\. 47–60\.Cited by:[§1](https://arxiv.org/html/2605.13932#S1.p1.1)\.
- \[49\]Y\. Wang, T\. Wang, S\. Li, X\. He, M\. Li, Z\. Wang, N\. Zheng, B\. Shao, and T\. Liu\(2024\)Enhancing geometric representations for molecules with equivariant vector\-scalar interactive message passing\.Nature Communications15\(1\),pp\. 313\.Cited by:[Table 3](https://arxiv.org/html/2605.13932#A2.T3.9.9.10.1.3),[§1](https://arxiv.org/html/2605.13932#S1.p1.1),[§2](https://arxiv.org/html/2605.13932#S2.p1.1),[Table 1](https://arxiv.org/html/2605.13932#S3.T1.9.9.10.1.2)\.
- \[50\]Z\. Wang, Z\. Dai, B\. Póczos, and J\. Carbonell\(2019\)Characterizing and avoiding negative transfer\.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition,pp\. 11293–11302\.Cited by:[§2](https://arxiv.org/html/2605.13932#S2.p3.1)\.
- \[51\]Z\. Wu, B\. Ramsundar, E\. N\. Feinberg, J\. Gomes, C\. Geniesse, A\. S\. Pappu, K\. Leswing, and V\. Pande\(2018\)MoleculeNet: a benchmark for molecular machine learning\.Chemical science9\(2\),pp\. 513–530\.Cited by:[§1](https://arxiv.org/html/2605.13932#S1.p2.1),[§2](https://arxiv.org/html/2605.13932#S2.p2.1)\.
- \[52\]H\. Xiang, J\. Xia, X\. Jin, W\. Du, L\. Zeng, and X\. Zeng\(2025\)Electron density\-enhanced molecular geometry learning\.InProceedings of the Thirty\-Fourth International Joint Conference on Artificial Intelligence,pp\. 7840–7848\.Cited by:[§2](https://arxiv.org/html/2605.13932#S2.p1.1)\.
- \[53\]N\. Yang, K\. Zeng, Q\. Wu, X\. Jia, and J\. Yan\(2022\)Learning substructure invariance for out\-of\-distribution molecular representations\.Advances in Neural Information Processing Systems35,pp\. 12964–12978\.Cited by:[§1](https://arxiv.org/html/2605.13932#S1.p2.1),[§2](https://arxiv.org/html/2605.13932#S2.p2.1)\.
- \[54\]J\. Yu, Z\. Wu, J\. Cai, A\. L\. Jia, and J\. Fan\(2024\)Kernel readout for graph neural networks\.\.InIJCAI,pp\. 2505–2514\.Cited by:[§2](https://arxiv.org/html/2605.13932#S2.p1.1)\.
- \[55\]J\. Yu, Z\. Wu, J\. Lu, T\. Wang, and H\. Wang\(2025\)A centrality\-based graph learning framework\.InProceedings of the Thirty\-Fourth International Joint Conference on Artificial Intelligence,pp\. 3588–3596\.Cited by:[§2](https://arxiv.org/html/2605.13932#S2.p1.1)\.
- \[56\]H\. Zhao, S\. Zhang, G\. Wu, J\. M\. Moura, J\. P\. Costeira, and G\. J\. Gordon\(2018\)Adversarial multiple source domain adaptation\.Advances in neural information processing systems31\.Cited by:[§2](https://arxiv.org/html/2605.13932#S2.p3.1)\.
- \[57\]Z\. Zhou, S\. Kearnes, L\. Li, R\. N\. Zare, and P\. Riley\(2019\)Optimization of molecules via deep reinforcement learning\.Scientific reports9\(1\),pp\. 10752\.Cited by:[§1](https://arxiv.org/html/2605.13932#S1.p1.1)\.

## Appendix ATraining Algorithm and Implementation Details

All experiments are conducted on NVIDIA RTX A6000 \(48 GB\) GPUs\. The computational overhead ofPOMAcan be decoupled into two phases: \(1\) A one\-time offline GRPO policy optimization, requiring approximately 72 GPU hours to converge; and \(2\) an online target adaptation phase, which leverages the pre\-trained policy for source selection and dual\-scale alignment\. Under the default configuration \(M=50M=50,K=5K=5, 40 GRPO steps\), the online phase requires only about 4\.5 GPU hours per target scaffold, demonstrating high deployment efficiency for practical screening scenarios\.

Algorithm 1POMA: Target\-Aware Source Selection and Adaptation1:Source pool

𝒮p​o​o​l\\mathcal\{S\}\_\{pool\}, unlabeled targets

𝒯r​e​a​l\\mathcal\{T\}\_\{real\}, backbone

fθf\_\{\\theta\}, policy

πθ\\pi\_\{\\theta\}
2:Adapted model for zero\-shot extrapolation on

𝒯r​e​a​l\\mathcal\{T\}\_\{real\}
3:// Stage 1: TSEC — Proxy Target Construction

4:Compute

HubScore⁡\(s\)\\operatorname\{HubScore\}\(s\)for each

s∈𝒮p​o​o​ls\\in\\mathcal\{S\}\_\{pool\}via Equation \([3](https://arxiv.org/html/2605.13932#S3.E3)\)

5:Select top\-

NpN\_\{p\}scaffolds as proxy targets

𝒯p​r​o​x​y\\mathcal\{T\}\_\{proxy\}
6:Rank remaining candidates via

Sr​a​n​kS\_\{rank\}; retain top\-

MMas candidate pool

𝒞\\mathcal\{C\}
7:// Stage 2: GRPO — Policy Optimization

8:foreach GRPO iterationdo

9:Construct state matrix

𝐗=\[𝐱1,…,𝐱M\]\\mathbf\{X\}=\[\\mathbf\{x\}\_\{1\},\\dots,\\mathbf\{x\}\_\{M\}\]\(Equation \([7](https://arxiv.org/html/2605.13932#S3.E7)\)\)

10:Sample

GGaction vectors

\{𝐚0,𝐚1,…,𝐚G−1\}\\\{\\mathbf\{a\}\_\{0\},\\mathbf\{a\}\_\{1\},\\dots,\\mathbf\{a\}\_\{G\-1\}\\\}
11:foreach valid action

𝐚i\\mathbf\{a\}\_\{i\}do

12:Execute dual\-scale adaptation on

𝒯p​r​o​x​y\\mathcal\{T\}\_\{proxy\}\(Equation \([5](https://arxiv.org/html/2605.13932#S3.E5)\)\)

13:Compute reward

Ri=MAEb​a​s​e−MAEiR\_\{i\}=\\mathrm\{MAE\}\_\{base\}\-\\mathrm\{MAE\}\_\{i\}
14:endfor

15:Compute advantages

A^i\\hat\{A\}\_\{i\}\(Equation \([9](https://arxiv.org/html/2605.13932#S3.E9)\)\)

16:Update

πθ\\pi\_\{\\theta\}via

ℒG​R​P​O\\mathcal\{L\}\_\{GRPO\}\(Equation \([11](https://arxiv.org/html/2605.13932#S3.E11)\)\)

17:endfor

18:// Stage 3: Inference — Zero\-Shot Adaptation

19:Select

𝒮∗←πθ​\(𝒞\)\\mathcal\{S\}^\{\*\}\\leftarrow\\pi\_\{\\theta\}\(\\mathcal\{C\}\)\(single forward pass\)

20:Fine\-tune

fθf\_\{\\theta\}on

𝒮∗\\mathcal\{S\}^\{\*\}with

ℒt​o​t​a​l\\mathcal\{L\}\_\{total\}\(Equation \([6](https://arxiv.org/html/2605.13932#S3.E6)\)\)

21:returnAdapted

fθf\_\{\\theta\}

## Appendix BDetailed Performance Analysis across Individual Scaffolds

To provide a comprehensive evaluation of the robustness ofPOMA, we present the detailed mean absolute error for each of the fifteen target scaffolds identified in theSCOPE\-Benchprotocol\. Table[3](https://arxiv.org/html/2605.13932#A2.T3)summarizes these results across three distinct equivariant architectures and three fundamental molecular properties\. This granular breakdown is essential because the extreme out\-of\-distribution shifts in our benchmark are highly scaffold\-dependent\. The empirical evidence demonstrates thatPOMAconsistently outperforms the supervised baseline in the vast majority of testing scenarios\. For instance, on the ETNN architecture, our framework achieves superior performance in over eighty percent of the individual scaffold tasks\. This consistency indicates that the reinforcement learning policy successfully identifies structural priors that remain invariant even when the target scaffold is significantly different from the training distribution\. We also observe that the performance gains are particularly pronounced in scaffolds with high structural complexity such as Scaffold 7 and Scaffold 15\. In these cases, conventional fine\-tuning often leads to catastrophic forgetting or negative transfer due to the injection of irrelevant source noise\. By contrast, the target\-aware selection mechanism inPOMAisolates a sparse but highly relevant subset of source knowledge\. Consequently, the model maintains high predictive accuracy without being compromised by the topological gap\. The stability of these results across fifteen independent experiments further confirms that our policy optimization approach is not sensitive to the specific choice of proxy targets\. Instead, the framework learns a generalized selection logic that effectively bridges the gap between heterogeneous molecular domains\. This detailed evidence supports the conclusion that target\-aware source selection is a necessary prerequisite for reliable molecular property prediction in real\-world discovery pipelines\.

Table 3:Detailed mean absolute error \(MAE\) on each of the 15 individual target scaffolds ofSCOPE\-Bench\. For each scaffold and property, the better performance between the supervised baseline andPOMAis highlighted in bold\.
## Appendix CGranular Evaluation of Knowledge Extraction and Module Synergy

To investigate the contribution of each component to the overall performance, Table[4](https://arxiv.org/html/2605.13932#A3.T4)provides a granular mean absolute error analysis across fifteen distinct target scaffolds using various retrieval and alignment configurations\. This extensive breakdown reveals that heuristic source selection methods lack the necessary precision to handle the diverse topological landscapes of theSCOPE\-Benchprotocol\. While methods based on physical descriptors or graph kernels occasionally yield improvements on specific scaffolds, their performance remains inconsistent across the entire benchmark\. This instability suggests that static similarity metrics are insufficient for capturing the complex structural synergies required for optimal molecular property prediction\.

The results further underscore the necessity of our dual\-scale alignment architecture\. Eliminating either the macroscopic topological alignment or the microscopic pharmacophore alignment leads to a significant reduction in predictive precision across multiple targets\. Specifically, the removal of global covariance alignment hinders the ability of the model to capture large\-scale molecular trends, while the absence of sub\-structure alignment prevents the extraction of local chemical invariants\. The fullPOMAframework successfully integrates these two scales, ensuring that the transferred knowledge is both structurally relevant and task\-specific\.

The primary advantage of our reinforcement learning policy lies in its ability to navigate the exponentially large candidate pool to identify the most synergistic source combination for any given target\. Rather than relying on rigid similarity thresholds, the policy learns a dynamic selection logic that maximizes the extraction of informative structural priors\. This leads to a substantial performance boost on challenging targets such as Scaffold 7 and Scaffold 15, where the structural gap is most pronounced\. By effectively leveraging the most valuable source domains,POMAestablishes an optimal performance ceiling that consistently exceeds the capabilities of both pure supervised learning and naive adaptation strategies\.

Table 4:Detailed ablation study results measured in mean absolute error across fifteen individual target scaffolds for the HOMO property using the ViSNet architecture\. The best performance for each target is highlighted in bold\. The full framework is denoted asPOMA\.
## Appendix DGranular Sensitivity Analysis of Hyperparameters

To supplement the aggregated sensitivity trends presented in the main text, Table[5](https://arxiv.org/html/2605.13932#A4.T5)provides the complete mean absolute error distribution across all fifteen target scaffolds under various hyperparameter configurations\. This granular perspective is critical for understanding how the reinforcement learning policy adapts to different architectural constraints\.

The empirical results demonstrate that while extreme hyperparameter settings occasionally yield marginal improvements on isolated targets, they fail to provide consistent regularization across the diverse chemical landscape\. For instance, reducing the candidate pool size to twenty restricts the exploratory space of the policy\. This restriction causes a performance deterioration on complex structures, such as Scaffold 7 and Scaffold 15, because the policy is forced to select from suboptimal source combinations\. Conversely, expanding the pool size to one hundred dilutes the reward signal and introduces low\-quality topological noise\.

A similar pattern emerges when analyzing the subset size\. Restricting the selection to a single source domain fails to capture the necessary complementary structural invariants required for robust out\-of\-distribution generalization\. Although a subset size of three shows competitive performance on specific subsets, it lacks the critical mass needed to bridge extreme distribution gaps effectively\. The adopted default configuration robustly achieves the lowest prediction error on the vast majority of target scaffolds\. This extensive validation confirms that our selected hyperparameter configuration establishes the optimal balance between source diversity and selection precision without relying on dataset\-specific tuning\.

Table 5:Detailed hyperparameter sensitivity analysis measured in mean absolute error across fifteen individual target scaffolds for the HOMO property using the ViSNet architecture\. The best performance for each target is highlighted in bold\. The default configuration represents the optimal balance adopted throughout all main experiments\.
## Appendix EHyperparameter Settings

Table[6](https://arxiv.org/html/2605.13932#A5.T6)summarizes all key hyperparameters used across experiments\.

Table 6:Hyperparameter settings forPOMA\.

Similar Articles

Beyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) Benchmark

Hugging Face Daily Papers

The Nanotechnology Molecular Optimization (NMO) Benchmark introduces physics-based molecular design tasks replacing drug-discovery-focused metrics, aiming to drive scientific discovery in nanotechnology. The paper shows that advanced methods underperform simpler approaches on NMO, and proposes new baseline methods including a novel representation and domain-agnostic pretraining.

Controllable Molecular Generative Foundation Models

arXiv cs.LG

Proposes CoMole, a controllable molecular generative foundation model using motif-aware graph diffusion and reinforcement learning, achieving superior controllability across materials and drug discovery benchmarks.