利用基于深度学习的形态学轮廓改进分子形态对比预训练
摘要
本文通过使用基于深度学习的形态学轮廓来扩展MoCoP方法,以改善药物发现中用于QSAR预测和毒性评估的分子嵌入。
arXiv:2609.30433v1 Announce Type: new
Abstract: Recent advancements in image-based profiling techniques have enabled the collection of high-volume cell morphology data, allowing new molecular embedding models to learn from the experimental phenotypic perturbations of a molecule in a cell. Previously, we developed Molecule-Morphology Contrastive Pretraining (MoCoP), a strategy for aligning small molecule embeddings to morphology fingerprints extracted through CellProfiler. The resulting molecular representation showed transferable performance for quantitative structure--activity relationship (QSAR) prediction tasks. Here, we extend the method by using a deep-learning-based cell image encoding pipeline to extract more feature-rich morphology profiles and align them to the molecular embeddings through contrastive learning. The new embeddings encode more accurate information on how molecules perturb cell morphology and enable improvements for QSAR predictions through either fixed-embedding linear probes or fully flexible fine-tuning. Morphology retrieval performance scales log-linearly with training data size, suggesting continued improvements as larger datasets become available. The improved MoCoP v2 also achieves superior performance on toxicity prediction and competitive results on ADME and activity benchmarks, when compared with existing molecular embedding models that use both cell morphology and transcriptomic data during training.
查看缓存全文
缓存时间: 2026/09/29 09:36
# Improving Molecular-Morphology Contrastive Pretrainingusing Deep-Learning-based Morphology Profiles
Source: [https://arxiv.org/html/2609.30433](https://arxiv.org/html/2609.30433)
Jie Li††thanks:Corresponding author: jerry\.8\.li@gsk\.comKathryn E\. KirchoffAffiliation:GSK, CheminformaticsDante A\. PertusiAffiliation:GSK, CheminformaticsZhizhuo ZhangAffiliation:GSK, Artificial Intelligence and Machine Learning
###### Abstract
Recent advancements in image\-based profiling techniques have enabled the collection of high\-volume cell morphology data, allowing new molecular embedding models to learn from the experimental phenotypic perturbations of a molecule in a cell\. Previously, we developed Molecule\-Morphology Contrastive Pretraining \(MoCoP\), a strategy for aligning small molecule embeddings to morphology fingerprints extracted through CellProfiler\. The resulting molecular representation showed transferable performance for quantitative structure–activity relationship \(QSAR\) prediction tasks\. Here, we extend the method by using a deep\-learning\-based cell image encoding pipeline to extract more feature\-rich morphology profiles and align them to the molecular embeddings through contrastive learning\. The new embeddings encode more accurate information on how molecules perturb cell morphology and enable improvements for QSAR predictions through either fixed\-embedding linear probes or fully flexible fine\-tuning\. Morphology retrieval performance scales log\-linearly with training data size, suggesting continued improvements as larger datasets become available\. The improvedMoCoPv2 also achieves superior performance on toxicity prediction and competitive results on ADME and activity benchmarks, when compared with existing molecular embedding models that use both cell morphology and transcriptomic data during training\.
## Introduction
The identification of molecular representations that encode biologically relevant information is a central challenge in computational drug discovery\. Traditional molecular descriptors such as extended\-connectivity fingerprints \(ECFPs\)\[[44](https://arxiv.org/html/2609.30433#bib.bib47)\]capture substructural motifs but are inherently limited to information derivable from chemical graphs alone\. Recent progress in deep learning has produced a rich landscape of learned molecular representations, including graph neural networks \(GNNs\) pretrained with self\-supervised objectives\[[23](https://arxiv.org/html/2609.30433#bib.bib25),[45](https://arxiv.org/html/2609.30433#bib.bib26),[32](https://arxiv.org/html/2609.30433#bib.bib34),[61](https://arxiv.org/html/2609.30433#bib.bib33),[52](https://arxiv.org/html/2609.30433#bib.bib31),[58](https://arxiv.org/html/2609.30433#bib.bib32)\]and transformer\-based models operating on SMILES strings\[[46](https://arxiv.org/html/2609.30433#bib.bib28),[1](https://arxiv.org/html/2609.30433#bib.bib29),[10](https://arxiv.org/html/2609.30433#bib.bib30)\]\. Three\-dimensional molecular structure has also been incorporated through frameworks such as Uni\-Mol\[[62](https://arxiv.org/html/2609.30433#bib.bib27)\], Newtonnet\[[19](https://arxiv.org/html/2609.30433#bib.bib61)\]and PaiNN\[[49](https://arxiv.org/html/2609.30433#bib.bib62)\]\. While these methods learn expressive representations from chemical data, they do not directly capture how a molecule interacts with the complex machinery of a living cell\.
In parallel, image\-based profiling, particularly the Cell Painting assay\[[4](https://arxiv.org/html/2609.30433#bib.bib4),[17](https://arxiv.org/html/2609.30433#bib.bib5)\], has emerged as a scalable technique for measuring cellular responses to chemical perturbations\. By staining six cellular compartments and acquiring multiplexed fluorescence images, Cell Painting yields high\-dimensional morphological readouts that capture drug\-induced phenotypic changes\[[59](https://arxiv.org/html/2609.30433#bib.bib40),[60](https://arxiv.org/html/2609.30433#bib.bib39)\]\. Large\-scale joint efforts such as the Joint Undertaking for Morphological Profiling \(JUMP\) Cell Painting Consortium have generated image datasets spanning over 116,000 compounds\[[5](https://arxiv.org/html/2609.30433#bib.bib2),[6](https://arxiv.org/html/2609.30433#bib.bib3)\], creating an unprecedented resource for studying chemical biology at scale\. Several studies have demonstrated the utility of these morphological profiles for predicting compound bioactivity\[[38](https://arxiv.org/html/2609.30433#bib.bib35),[22](https://arxiv.org/html/2609.30433#bib.bib36),[50](https://arxiv.org/html/2609.30433#bib.bib37),[20](https://arxiv.org/html/2609.30433#bib.bib42)\], mechanism of action\[[55](https://arxiv.org/html/2609.30433#bib.bib38),[54](https://arxiv.org/html/2609.30433#bib.bib41)\], and drug targets\[[24](https://arxiv.org/html/2609.30433#bib.bib43)\]\.
Motivated by the complementary nature of chemical structure and cellular morphology, a growing body of work has explored cross\-modal contrastive learning to align molecular and phenotypic representations\. CLOOME\[[47](https://arxiv.org/html/2609.30433#bib.bib10)\]adapted the CLIP framework\[[41](https://arxiv.org/html/2609.30433#bib.bib46)\]to learn joint embeddings of chemical structures and Cell Painting images\. More recently, InfoAlign\[[31](https://arxiv.org/html/2609.30433#bib.bib8)\]introduced an information\-theoretic approach that integrates molecules with both morphological and transcriptomic data through a context graph, achieving superior performance on multiple downstream benchmarks\. CHMR\[[28](https://arxiv.org/html/2609.30433#bib.bib9)\]further improved upon InfoAlign by modeling hierarchical dependencies across molecular, cellular, and genomic levels, achieveing state\-of\-the\-art performance on multiple biological assay readout predictions, and absorption, distribution, metabolism, excretion and toxicity \(ADMET\) prediction\. Other notable approaches include MolPhenix\[[15](https://arxiv.org/html/2609.30433#bib.bib11)\], which demonstrated improved retrieval through pre\-trained phenomics models; PhenoScreen\[[57](https://arxiv.org/html/2609.30433#bib.bib13)\], which employed dual\-space contrastive learning; MINER\[[42](https://arxiv.org/html/2609.30433#bib.bib12)\], which addressed negative sampling calibration; and CellCLIP\[[35](https://arxiv.org/html/2609.30433#bib.bib14)\], which leveraged text\-guided contrastive learning\. Cross\-modal strategies involving transcriptomics have also been explored\[[18](https://arxiv.org/html/2609.30433#bib.bib16),[3](https://arxiv.org/html/2609.30433#bib.bib15),[14](https://arxiv.org/html/2609.30433#bib.bib49)\]\.
Despite these advances, a critical bottleneck in molecule–morphology alignment lies in the quality of the morphological representations themselves\. The conventional approach relies on CellProfiler\[[37](https://arxiv.org/html/2609.30433#bib.bib6)\], a rule\-based image analysis pipeline that extracts handcrafted features such as intensity statistics, texture, and shape descriptors\. While CellProfiler features are interpretable and widely used, they may miss subtle morphological patterns that are difficult to capture with predefined feature extractors\. Recent work in self\-supervised representation learning for microscopy has shown that deep learning models, including DINO\[[8](https://arxiv.org/html/2609.30433#bib.bib19),[7](https://arxiv.org/html/2609.30433#bib.bib20)\], masked autoencoders\[[26](https://arxiv.org/html/2609.30433#bib.bib18),[27](https://arxiv.org/html/2609.30433#bib.bib51)\], and supervised models\[[39](https://arxiv.org/html/2609.30433#bib.bib21),[25](https://arxiv.org/html/2609.30433#bib.bib17)\], can learn cellular representations that outperform CellProfiler on downstream tasks such as mechanism\-of\-action prediction and batch effect correction\[[2](https://arxiv.org/html/2609.30433#bib.bib22),[30](https://arxiv.org/html/2609.30433#bib.bib23)\]\. This suggests that replacing CellProfiler features with learned morphological embeddings could substantially improve molecule–morphology contrastive pretraining\.
Previously, we developed Molecule\-Morphology Contrastive Pretraining \(MoCoP\)\[[40](https://arxiv.org/html/2609.30433#bib.bib1)\], which aligned molecular GNN embeddings with CellProfiler\-derived morphological profiles using data from the JUMP\-CP Consortium\.MoCoPv1 demonstrated that the resulting molecular representations consistently improved GNN performance on quantitative structure–activity relationship \(QSAR\) prediction tasks across varying dataset sizes\. In this work, we introduceMoCoPv2, which replaces CellProfiler features with deep\-learning \(DL\) morphological embeddings as the alignment target\. These DL embeddings are extracted from Cell Painting images of the JUMP\-CP Consortium using a fine\-tuned convolutional neural network, followed by CORAL\-based batch correction\[[51](https://arxiv.org/html/2609.30433#bib.bib7)\]to mitigate inter\-source variability\. We show that this upgrade produces molecular representations that \(1\) more accurately reflect how molecules perturb cell morphology, \(2\) improve QSAR prediction on ChEMBL20 benchmarks, \(3\) exhibit log\-linear scaling of retrieval performance with training data size, and \(4\) achieve competitive or state\-of\-the\-art performance on downstream toxicity and ADME prediction tasks compared to recent methods including InfoAlign and CHMR, despite using simpler architecture, and only using cell morphology without transcriptomic data\.
## Methods
TheMoCoPv2 framework consists of three stages \(Figure[1](https://arxiv.org/html/2609.30433#Sx2.F1)\): \(1\) training a deep\-learning image encoder on Cell Painting images to extract DL morphological embeddings, including shading correction, tile extraction, and CORAL batch correction; \(2\) molecule–morphology contrastive pretraining, in which a molecular GNN encoder is aligned to the DL embedding space while the image encoder remains frozen; and \(3\) transfer of the pretrained molecular encoder to downstream applications\.
Figure 1:Overview of theMoCoPv2 pipeline\.\(a\)DL embedding extraction: raw 5\-channel Cell Painting images undergo shading correction and tile extraction \(224×224224\\times 224px\), followed by training and featurization with a ResNet\-18 model fine\-tuned using triplet loss\. The resulting tile\-level embeddings are mean\-aggregated per well to produce 128\-dimensional DL morphological embeddings, which are then batch\-corrected using CORAL alignment based on DMSO negative controls\.\(b\)Contrastive pretraining: batch\-corrected DL embeddings are encoded by a morphology multi\-layer perceptron network \(MLP\), while molecular graphs are encoded by a gated graph neural network \(GGNN\)\. Both branches are projected into a shared 128\-dimensional space, and the encoders are jointly trained with a symmetric InfoNCE contrastive loss\.\(c\)Downstream applications: the pretrained GGNN molecular encoder is transferred to new molecules for QSAR prediction, ADMET prediction, molecular property prediction, mechanism\-of\-action retrieval, and other tasks\.### Deep\-Learning\-based Morphological Embeddings
The primary methodological advance inMoCoPv2 is the replacement of CellProfiler\-derived morphological features with deep\-learning\-based \(DL\) embeddings\. WhereasMoCoPv1 used CellProfiler to extract handcrafted features \(intensity, texture, shape descriptors\) from Cell Painting images, DL embeddings are learned representations obtained by fine\-tuning a convolutional neural network to distinguish compound treatments from their morphological effects\. The DL embedding extraction pipeline proceeds in three stages: image preprocessing, model training, and inference\.
Image Preprocessing\.Raw Cell Painting images are first corrected for uneven illumination using an illumination correction function \(ICF\) estimated per channel for each plate\. The ICF is computed by smoothing the 10th\-percentile intensity image with a Gaussian filter \(σ=50\\sigma=50, kernel size=250=250pixels\), and each raw image is divided by this correction field\. Next, individual cells are localized using a Laplacian\-of\-Gaussian \(LoG\) nuclei detector with sigma parameters scaled by the microscope magnification \(at20×20\\times:σmin=10\\sigma\_\{\\min\}=10,σmax=20\\sigma\_\{\\max\}=20pixels\), and a224×224224\\times 224pixel tile is extracted centered on each detected nucleus, with overlapping tiles filtered to avoid redundancy\. All five fluorescence channels \(DNA, endoplasmic reticulum, actin/Golgi/plasma membrane, mitochondria, and nucleoli/RNA\) are bundled into each tile and stored as HDF5 files for efficient data loading\.
Model Architecture and Training\.We employed a ResNet\-18\[[21](https://arxiv.org/html/2609.30433#bib.bib44)\]backbone pretrained on ImageNet\. Since the standard model expects three\-channel RGB input, the first convolutional layer \(conv1\) was replaced with a new layer accepting five input channels while preserving the original 64 output feature maps,7×77\\times 7kernel size, stride of 2, and padding of 3; pretrained weights for all subsequent layers were retained\. The backbone feeds into an MLP projection head with hidden dimensions of 1024 and 128, producing 128\-dimensional tile\-level embeddings\. During both training and inference, 12 tiles are randomly sampled per well, passed through the network independently, and mean\-aggregated to produce a single well\-level DL embedding\.
The model was fine\-tuned using a triplet margin loss with online semihard mining\[[48](https://arxiv.org/html/2609.30433#bib.bib50)\]\. For each valid triplet\(za,zp,zn\)\(z\_\{a\},z\_\{p\},z\_\{n\}\), wherezaz\_\{a\}is an anchor,zpz\_\{p\}a positive \(same compound\), andznz\_\{n\}a negative \(different compound\), the per\-triplet loss is:
ℓ\(za,zp,zn\)=max\(0,‖za−zp‖2−‖za−zn‖2\+m\)\\ell\(z\_\{a\},z\_\{p\},z\_\{n\}\)=\\max\\bigl\(0,\\;\\\|z\_\{a\}\-z\_\{p\}\\\|\_\{2\}\-\\\|z\_\{a\}\-z\_\{n\}\\\|\_\{2\}\+m\\bigr\)\(1\)wherem=0\.1m=0\.1is the margin\. Semihard mining retains only triplets where the negative is farther than the positive but within the margin \(0<ℓ<m0<\\ell<m\), and the final loss is the mean over all such triplets in the batch; embeddings are not L2\-normalized before distance computation\. A balanced sampler ensured each training batch contained 8 distinct compound classes with2×2\\timesoversampling to handle class imbalance\. The model was trained on all 217 plates from Source 3 of the JUMP\-CP Consortium using AdamW\[[34](https://arxiv.org/html/2609.30433#bib.bib53)\]for 100 epochs on 4 NVIDIA A100 80 GB GPUs, with 1\-nearest\-neighbor Euclidean accuracy as the primary evaluation metric\. Source 3 was selected for image encoder training because it is the largest single data source in the JUMP\-CP dataset, providing 54,655 wells across 217 plates with consistent imaging conditions from a single laboratory\. Training the image encoder on a single source avoids conflating morphological variation with batch effects during the supervised triplet learning stage; inter\-source variability is instead addressed downstream through CORAL batch correction\. More details about training can be found in[Appendix A](https://arxiv.org/html/2609.30433#A1)\.
Inference\.After training, the model was applied to extract DL embeddings for every well across all 1,728 plates in the full JUMP\-CP dataset, spanning multiple data sources and experimental batches\. Importantly, the image encoder is trained once and then frozen: the resulting DL embeddings are pre\-extracted and stored as fixed 128\-dimensional vectors\. During the subsequent molecule–morphology contrastive pretraining, the DL embeddings serve as static input to a trainable morphology projection network; the image encoder itself is not revisited\.
### Batch Correction with CORAL
Cell Painting experiments conducted across different laboratories, instruments, and time points introduce substantial batch effects that can dominate the embedding space, causing representations to cluster by data source rather than by biological perturbation\[[2](https://arxiv.org/html/2609.30433#bib.bib22),[53](https://arxiv.org/html/2609.30433#bib.bib24)\]\. To address this, we applied CORrelation ALignment \(CORAL\)\[[51](https://arxiv.org/html/2609.30433#bib.bib7)\]to correct batch effects in the extracted DL embeddings\. This batch correction step is essential: without it, the DL embeddings are dominated by data\-source identity rather than compound\-induced morphological changes \(see[Appendix C](https://arxiv.org/html/2609.30433#A3)for visualization\)\.
CORAL aligns the second\-order statistics \(covariance matrices\) of embeddings from different batches\. Specifically, letCsC\_\{s\}andCtC\_\{t\}denote the covariance matrices of DL embeddings for DMSO\-treated negative control wells in a source batch and the reference batch \(Source 3\), respectively\. The batch\-corrected embeddings for all compounds in the source batch are obtained by applying the transformation:
X^s=XsCs−1/2Ct1/2\\hat\{X\}\_\{s\}=X\_\{s\}\\,C\_\{s\}^\{\-1/2\}\\,C\_\{t\}^\{1/2\}\(2\)whereXsX\_\{s\}is the matrix of original embeddings from the source batch\. By aligning on the DMSO controls that exhibit minimal compound\-induced variation, the correction preserves biologically meaningful signals while removing technical variability\.
### Molecule\-Morphology Contrastive Pretraining
Following the framework introduced inMoCoPv1\[[40](https://arxiv.org/html/2609.30433#bib.bib1)\], we jointly learn a molecular encoder and a morphology encoder using contrastive learning on paired \(molecule, morphology\) data\. The pretraining dataset consists ofNNmolecule–morphology pairs\{\(ximol,ximorph\)∣i∈\{1,…,N\}\}\\\{\(x\_\{i\}^\{\\text\{mol\}\},x\_\{i\}^\{\\text\{morph\}\}\)\\mid i\\in\\\{1,\\ldots,N\\\}\\\}\.
Encoders and Projections\.The molecular encoderfmolf\_\{\\text\{mol\}\}is a Gated Graph Neural Network \(GGNN\)\[[29](https://arxiv.org/html/2609.30433#bib.bib45)\]that operates on 2D molecular graphs where atoms are nodes \(with 75\-dimensional features\) and bonds are edges\. The GGNN consists of 6 message\-passing layers followed by a fully connected layer of dimension 1024, with dropout \(p=0\.1p=0\.1\)\. The morphology encoderfmorphf\_\{\\text\{morph\}\}is an MLP with hidden layer dimensions \[512, 256, 128\] and dropout \(p=0\.1p=0\.1\), operating on the 128\-dimensional batch\-corrected DL embeddings\.
Each encoder produces a representation that is transformed via a projection functiongginto a shared embedding space:
himol\\displaystyle h\_\{i\}^\{\\text\{mol\}\}=fmol\(ximol\),uimol=normalize\(gmol\(σ\(himol\)\)\)\\displaystyle=f\_\{\\text\{mol\}\}\(x\_\{i\}^\{\\text\{mol\}\}\),\\quad u\_\{i\}^\{\\text\{mol\}\}=\\text\{normalize\}\\bigl\(g\_\{\\text\{mol\}\}\(\\sigma\(h\_\{i\}^\{\\text\{mol\}\}\)\)\\bigr\)\(3\)himorph\\displaystyle h\_\{i\}^\{\\text\{morph\}\}=fmorph\(ximorph\),uimorph=normalize\(gmorph\(σ\(himorph\)\)\)\\displaystyle=f\_\{\\text\{morph\}\}\(x\_\{i\}^\{\\text\{morph\}\}\),\\quad u\_\{i\}^\{\\text\{morph\}\}=\\text\{normalize\}\\bigl\(g\_\{\\text\{morph\}\}\(\\sigma\(h\_\{i\}^\{\\text\{morph\}\}\)\)\\bigr\)\(4\)whereσ\\sigmadenotes a ReLU non\-linearity,gmolg\_\{\\text\{mol\}\}andgmorphg\_\{\\text\{morph\}\}are learned linear projections toℝ128\\mathbb\{R\}^\{128\}, andnormalize\(⋅\)\\text\{normalize\}\(\\cdot\)denotes L2 normalization\.
Contrastive Objective\.The two encoders are jointly optimized using a symmetric InfoNCE loss\[[56](https://arxiv.org/html/2609.30433#bib.bib56),[41](https://arxiv.org/html/2609.30433#bib.bib46)\]\. Within a mini\-batch ofNNpairs, the cosine similarity matrixSij=⟨uimol,ujmorph⟩S\_\{ij\}=\\langle u\_\{i\}^\{\\text\{mol\}\},u\_\{j\}^\{\\text\{morph\}\}\\rangleis computed, and the loss is defined as:
ℒcontrastive=12\(ℒmol→morph\+ℒmorph→mol\)\\mathcal\{L\}\_\{\\text\{contrastive\}\}=\\frac\{1\}\{2\}\\bigl\(\\mathcal\{L\}\_\{\\text\{mol\}\\to\\text\{morph\}\}\+\\mathcal\{L\}\_\{\\text\{morph\}\\to\\text\{mol\}\}\\bigr\)\(5\)where the directional losses are:
ℒmol→morph\\displaystyle\\mathcal\{L\}\_\{\\text\{mol\}\\to\\text\{morph\}\}=−1N∑i=1Nlogexp\(τ⋅Sii\)∑k=1Nexp\(τ⋅Sik\)\\displaystyle=\-\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\log\\frac\{\\exp\(\\tau\\cdot S\_\{ii\}\)\}\{\\sum\_\{k=1\}^\{N\}\\exp\(\\tau\\cdot S\_\{ik\}\)\}\(6\)ℒmorph→mol\\displaystyle\\mathcal\{L\}\_\{\\text\{morph\}\\to\\text\{mol\}\}=−1N∑i=1Nlogexp\(τ⋅Sii\)∑k=1Nexp\(τ⋅Ski\)\\displaystyle=\-\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\log\\frac\{\\exp\(\\tau\\cdot S\_\{ii\}\)\}\{\\sum\_\{k=1\}^\{N\}\\exp\(\\tau\\cdot S\_\{ki\}\)\}\(7\)withτ=10\\tau=10as a temperature scaling parameter applied during training\.
After pretraining, the molecular encoderfmolf\_\{\\text\{mol\}\}is transferred to downstream tasks, while the morphology encoder and projection heads are discarded\.
Training Details\.The model was trained with a batch size of 256 using AdamW\[[34](https://arxiv.org/html/2609.30433#bib.bib53)\]with cosine annealing warm restarts\[[33](https://arxiv.org/html/2609.30433#bib.bib55)\]\. Early stopping was applied based on validation retrieval accuracy with a patience of 500 epochs\. The maximum training duration was set to 1,000 epochs\. Canonical SMILES representations were used for molecular identity resolution, and all training was conducted using PyTorch Lightning\[[11](https://arxiv.org/html/2609.30433#bib.bib54)\]\.[Appendix D](https://arxiv.org/html/2609.30433#A4)describes additional details about training configurations\.
## Results and Discussion
### Morphology Retrieval Performance and Scaling Behavior
We first evaluated the morphology retrieval performance ofMoCoPv2\. In this evaluation, each test compound’s molecular graph is passed through the trained molecular encoder to produce a predicted embedding\. This predicted embedding is then compared against the morphological profiles of all compounds in a retrieval pool \(a batch of either 100 or 1,000 test compounds\), and the goal is to identify the correct morphological profile, i\.e\., the one belonging to the same compound\. Performance is measured by top\-kkaccuracy: the fraction of queries for which the correct match appears among thekknearest neighbors ranked by cosine similarity in the shared embedding space\. Higher retrieval accuracy indicates that the molecular encoder has learned to map chemical structures to points in embedding space that are close to their corresponding morphological profiles\.
To investigate how retrieval performance scales with training data, we trainedMoCoPv2 models using four subsets of the JUMP\-CP dataset: 10K, 20K, 50K, and 98K \(full\) training compounds\. All models were trained from scratch with identical hyperparameters, using the same DL morphological embeddings and test split\.
A clear scaling law is observed \(Figure[2](https://arxiv.org/html/2609.30433#Sx3.F2)\): retrieval accuracy improves consistently as more compounds are used for contrastive pretraining\. In the 1:100 retrieval setting, Top\-1 accuracy increases from 3\.3% \(10K compounds\) to 6\.8% \(98K compounds\), a2\.1×2\.1\\timesimprovement\. Top\-10 accuracy also rises from 22\.4% to 41\.6%, a1\.9×1\.9\\timesimprovement\. The 1:1000 retrieval setting exhibits a similar pattern, with Top\-1 accuracy improving from 0\.6% to 1\.6% \(2\.7×2\.7\\times\) and top\-10 accuracy improving from 4\.9% to 9\.8% \(2×2\\times\)\. The approximately log\-linear relationship between training data size and retrieval performance provides confidence that the framework can continue to benefit from additional data as larger Cell Painting datasets become available\.
Figure 2:Scaling behavior ofMoCoPv2 morphology retrieval\. Models are trained from scratch with 10K, 20K, 50K, and 98K compounds and evaluated on the same held\-out test set\.\(a\)1:100 retrieval \(100\-candidate pool\)\.\(b\)1:1000 retrieval \(1,000\-candidate pool\)\. Retrieval accuracy improves consistently with training data size across all top\-kkmetrics, following an approximately log\-linear trend\.
### MoCoP v2 Captures More Accurate Morphological Relationships
Next, we compared theMoCoPv1 and v2 embeddings in terms of how well they capture the morphological changes a compound introduces to the cells, and foundMoCoPv2 embeddings encode more biologically meaningful information about molecular perturbations thanMoCoPv1 embeddings through both qualitative and quantitative evaluations\. Pairwise cosine similarities betweenMoCoPv1 and v2 embeddings are strongly correlated \(Pearsonr=0\.873r=0\.873, Spearmanρ=0\.871\\rho=0\.871\), indicating that the two models agree on most compound pairs\. However, the approximately 13% of unexplained variance reveals cases where the models diverge, and in these cases, visual inspection of Cell Painting images consistently supports v2’s assessment over v1’s\.
Divergence Analysis: comparing raw Cell Painting images\.To understand the differences between v1 and v2 embeddings, we identified compound pairs where the two models disagree most strongly about morphological similarity from the test dataset\. For each pair, we computed the cosine similarity in both theMoCoPv1 and v2 embedding spaces after applying mean\-centering to address similarity collapse, and selected pairs with the largest difference in similarity between the two models\. All image comparisons were restricted to Source 3 compounds not used for MoCoP training, eliminating confounding batch effects and keeping comparisons fair\.
Figure[3](https://arxiv.org/html/2609.30433#Sx3.F3)a shows a representative case whereMoCoPv1 incorrectly assigns high similarity to two compounds that have visibly different phenotypes\. v1 assigns cosine similarity of\+0\.619\+0\.619while v2 assigns−0\.387\-0\.387\. The Cell Painting images reveal clear differences: Compound A shows larger, well\-spread cells at moderate density, while Compound B produces compact cell clusters with smaller cell sizes and more visible disintegration\. v2 correctly identifies these as dissimilar\.
Conversely, Figure[3](https://arxiv.org/html/2609.30433#Sx3.F3)b shows a case whereMoCoPv2 correctly identifies phenotypic similarity that v1 misses\. Despite a low Tanimoto similarity of 0\.226, v2 assigns high similarity \(\+0\.615\+0\.615\) while v1 assigns low similarity \(−0\.286\-0\.286\)\. Inspection of the Cell Painting images confirms that both compounds produce concordant phenotypes \(similar cell density, spreading pattern, and balance of contents from various visible channels\), which is correctly captured by the DL\-embedding\-aligned v2 embeddings but missed by the CellProfiler\-aligned v1 embeddings\.
[Appendix E](https://arxiv.org/html/2609.30433#A5)shows additional examples of embedding divergence betweenMoCoPv1 and v2, along with their corresponding Cell Painting images\. In all cases, the phenotypes observed in the images are more consistent with the v2 embedding similarities, confirming that v2 provides a more accurate representation of the morphological perturbations that compounds introduce to cells\.
\(a\)v1 incorrectly assigns high similarity to compounds with different phenotypes\. Compound A shows large, well\-spread cells; Compound B shows compact clusters with smaller cells \(v1 cos=\+0\.619=\+0\.619, v2 cos=−0\.387=\-0\.387, Tanimoto=0\.098=0\.098\)\.
\(b\)v2 correctly identifies phenotypic similarity missed by v1\. Both compounds produce concordant phenotypes with similar cell density and channel balance \(v2 cos=\+0\.615=\+0\.615, v1 cos=−0\.286=\-0\.286, Tanimoto=0\.226=0\.226\)\.
Figure 3:Representative examples of strong divergence betweenMoCoPv1 and v2 embeddings\. Cell Painting composite: Blue = DNA, Green = AGP, Red = Mitochondria\. Chemical structures are shown alongside each compound\. Plate barcodes and well positions from JUMP\-CP are also marked for reference\.\(a\)v1 overestimates similarity for a pair with distinct phenotypes; v2 correctly identifies them as dissimilar\.\(b\)v2 recognizes shared phenotypic effects between structurally dissimilar compounds that v1 misses\.Quantitative Context: Split\-Half Replicate Concordance\.To rigorously evaluate embedding quality, we designed a split\-half replicate concordance analysis that measures how well each embedding’s nearest neighbors capture*reproducible*morphological similarity\. For each of the 5,265 test compounds with at least four replicate wells in the JUMP\-CP dataset, we randomly split the wells into two independent halves \(A and B\) and computed compound\-level CellProfiler profiles by averaging features within each half\. This yields two independent morphological similarity estimates for every compound pair\. We then established an*oracle ceiling*: the retrieval quality achievable if molecular embeddings perfectly predicted morphological similarity\. Specifically, for each compound we selected its top\-kkmost morphologically similar neighbors using half A profiles and evaluated those neighbors’ morphological similarity using the independent half B profiles\. This cross\-validated ceiling represents the maximum retrieval quality that is reproducible across independent replicate measurements; it accounts for measurement noise and batch effects that limit any model’s achievable performance\.
For each embedding method \(MoCoPv2,MoCoPv1, and ECFP4 with 1024\-bit Morgan fingerprints\), we retrieved the top\-kknearest neighbors in embedding space and measured the mean CellProfiler cosine similarity of those neighbors, averaged over both replicate halves\. Results are expressed as a percentage of the oracle ceiling to normalize for the inherent difficulty of the task at eachkk\(Figure[4](https://arxiv.org/html/2609.30433#Sx3.F4)a\)\. The entire procedure was repeated across 50 random replicate splits to obtain bootstrap confidence intervals; the resulting error bars are small, confirming that the observed differences are robust to the choice of replicate partition\. NeitherMoCoPv1 nor v2 was trained on the CellProfiler features used for evaluation, making this an independent assessment of embedding quality\.
ECFP4 achieves the highest retrieval quality at smallkk\(44\.6%44\.6\\%of ceiling atk=1k=1\), reflecting that structurally similar compounds often induce similar phenotypes\. However, ECFP4’s performance declines askkincreases \(38\.6%38\.6\\%atk=10k=10\) because it exhausts the pool of close structural analogs\.MoCoPv2 maintains more stable performance across allkkvalues \(38\.838\.8–42\.0%42\.0\\%of ceiling\) and overtakes ECFP4 atk≥10k\\geq 10, demonstrating its ability to identify phenotypically similar compounds beyond the reach of structural fingerprints\.MoCoPv2 also consistently outperforms v1 at everykkvalue tested, with the strongest advantage atk=1k=1, showing the v2 embeddings are more capable of accurately identifying phenotypically similar compounds with immediate neighbors in the embedding space\. This is especially notable given thatMoCoPv1 was explicitly aligned to CellProfiler profiles, yet v2, which was aligned to deep\-learning morphological embeddings, better predicts CellProfiler\-based similarity\.
The advantage of v2 is most profound for*chemically isolated*compounds \(maximum Tanimoto similarity to any other compound<0\.3<0\.3\)\. In this regime, where structural fingerprints provide limited information, MoCoP v2 outperforms both MoCoP v1 and ECFP4 \(Figure[4](https://arxiv.org/html/2609.30433#Sx3.F4)b\) significantly\. This is precisely where morphology\-informed embeddings provide the greatest value: identifying molecules with distinct scaffolds that nonetheless share similar phenotypic effects\. Indeed, even among compound pairs with Tanimoto similarity below 0\.2,MoCoPv2 can still identify shared or divergent phenotypic effects that v1 misses\.[Appendix E](https://arxiv.org/html/2609.30433#A5)presents visual examples of such structurally dissimilar pairs where v2 captures morphological relationships invisible to both v1 and chemical fingerprints, highlighting the value of DL\-embedding\-aligned molecular representations for scaffold hopping applications\.
\(a\)Top\-kkCellProfiler retrieval \(% of oracle ceiling\)
\(b\)Stratified by chemical isolation \(k=10k=10\)
Figure 4:Quantitative comparison ofMoCoPv2, v1, and ECFP4 embedding quality, measured by CellProfiler profile retrieval on 5,265 JUMP test compounds with≥4\\geq 4replicate wells\.\(a\)Retrieval quality expressed as percentage of the oracle ceiling \(see text\), averaged over 50 random replicate splits\. Error bars indicate bootstrap standard deviation\. ECFP4 dominates at smallkkbut declines as structural analogs are exhausted; v2 overtakes ECFP4 atk≥10k\\geq 10and consistently outperforms v1 at allkkvalues\.\(b\)v2 provides the best morphology retrieval for chemically isolated compounds \(Tanimoto<0\.3<0\.3\), exactly the regime where structure\-based methods have the least information\.
### Improved QSAR Prediction on ChEMBL20
We evaluated the transferability ofMoCoPv2 molecular embeddings on QSAR prediction tasks using the ChEMBL20 dataset\[[36](https://arxiv.org/html/2609.30433#bib.bib57)\], following the exact procedure as described in MoCoP v1\. ChEMBL20 dataset comprises a diverse set of measurements including ADME, toxicity, physicalchemical properties, binding and functional assays, consisting of 1,310 binary downstream tasks of about 450K compounds\. Two evaluation protocols were used for comparison: \(1\)*linear probing*, where the pretrained molecular encoder is frozen and a single linear layer is trained on top; and \(2\)*end\-to\-end fine\-tuning*, where the full molecular encoder is updated jointly with a task\-specific head, with molecular encoder weights initialized from the pretrained checkpoint\. Both protocols were evaluated across three data regimes \(5%, 25%, and 100% of the ChEMBL20 data\) to assess data efficiency\. Results are reported as mean±\\pmstandard deviation across 3 train/validation/test splits\.
Linear Probe Results\.Table[1](https://arxiv.org/html/2609.30433#Sx3.T1)shows thatMoCoPv2 consistently outperformsMoCoPv1 across all data proportions for both AUROC and AUPRC, with an improvement of 3\.5 percentage points in AUROC and 2\.4 in AUPRC at 5% data; 3\.1 and 2\.9 at 25% data; and 2\.5 and 2\.6 at 100% data\. These consistent gains confirm that the DL\-embedding\-aligned representations fromMoCoPv2 encode more informative molecular features than CellProfiler\-aligned embeddings, which could be easily exploited by a linear model from the frozen embeddings\.
Table 1:Linear probe QSAR prediction performance on ChEMBL20\. Mean±\\pmstandard deviation across 3 splits\. Bold indicates the better result\.End\-to\-End Fine\-Tuning Results\.When allowing the full molecular encoder to be fine\-tuned on the downstream QSAR tasks, the results become more nuanced \(Table[2](https://arxiv.org/html/2609.30433#Sx3.T2)\)\.MoCoPv2 slightly underperforms v1 in the low\-data regimes \(5% and 25% AUROC\) but surpasses v1 at 100% data, with improvements of 1\.3 percentage points in AUROC and 3\.2 in AUPRC\. v2 achieves higher AUPRC than v1 across all data proportions, indicating better precision–recall trade\-offs\. Even though there are some overlap in the confidence intervals, the consistent direction of improvement in AUPRC across all three data proportions, combined with the similar gains observed in the linear probe setting \(Table[1](https://arxiv.org/html/2609.30433#Sx3.T1)\), supports the conclusion that v2 embeddings encode richer information\. The fine\-tuning results indicate that this richer information is most effectively exploited when sufficient downstream data is available for adaptation\.
Table 2:End\-to\-end fine\-tuning QSAR prediction performance on ChEMBL20\. Mean±\\pmstandard deviation across 3 splits\. Bold indicates the better result\.
### Comparing with Other State\-of\-the\-Art Methods
To contextualize the performance ofMoCoPv2 against existing state\-of\-the\-art molecular embedding methods that may incorporate one or more modalities, including molecular structure, cell morphology and transcriptomics, we evaluated on four benchmark datasets spanning diverse prediction tasks: ToxCast \(toxicity classification: 8,576 compounds across 617 tasks\)\[[43](https://arxiv.org/html/2609.30433#bib.bib58)\], ChEMBL 2K \(activity classification: 2,355 compounds across 41 tasks\)\[[16](https://arxiv.org/html/2609.30433#bib.bib59)\], Broad 6K \(activity classification: 6,567 compounds across 32 tasks\)\[[38](https://arxiv.org/html/2609.30433#bib.bib35)\], and Biogen 3K \(ADME regression: 3,521 compounds across 6 tasks\)\[[12](https://arxiv.org/html/2609.30433#bib.bib60)\]\. These benchmarks were curated from the InfoAlign repository\[[31](https://arxiv.org/html/2609.30433#bib.bib8)\]and are widely used for evaluating various molecular representations\. Train, validation and test splits also follow the same scaffold\-based splitting as in\[[31](https://arxiv.org/html/2609.30433#bib.bib8)\], which allows fair comparison across methods\. All results are reported as 3\-seed mean±\\pmstandard deviation on the test set\. A comprehensive hyperparameter sweep was conducted forMoCoPv2 to identify optimal configurations for each dataset \(see[Appendix F](https://arxiv.org/html/2609.30433#A6)\)\.
Table[3](https://arxiv.org/html/2609.30433#Sx3.T3)comparesMoCoPv2 with several baseline and state\-of\-the\-art methods\. To facilitate interpretation, the table is organized into three groups based on the training data modalities\. The first group consists of*structure\-only*baselines \(an MLP trained on Morgan fingerprints\[[44](https://arxiv.org/html/2609.30433#bib.bib47)\], MolT5\[[9](https://arxiv.org/html/2609.30433#bib.bib52)\], and Uni\-Mol\[[62](https://arxiv.org/html/2609.30433#bib.bib27)\]\), which use only molecular structure during training\. The second group contains methods that use*structure and cell morphology only*: CLOOME\[[47](https://arxiv.org/html/2609.30433#bib.bib10)\], which aligns molecular graphs with Cell Painting images, andMoCoPv2\. The third group includes InfoAlign\[[31](https://arxiv.org/html/2609.30433#bib.bib8)\]and CHMR\[[28](https://arxiv.org/html/2609.30433#bib.bib9)\], which use*structure, cell morphology, and transcriptomics*during training\. This distinction is important for a fair comparison: InfoAlign constructs a context graph that jointly encodes molecular, morphological, and gene expression modalities, while CHMR models hierarchical dependencies across molecular, cellular, and genomic levels using vector quantization\. Both methods therefore have access to additional biological information \(transcriptomic readouts\) that is unavailable toMoCoPv2 during training\.
Table 3:Performance comparison on downstream benchmark tasks\. Classification tasks \(ChEMBL 2K, ToxCast, Broad 6K\) report average AUC \(%,↑\\uparrow\)\. Biogen 3K reports average MAE \(×100\\times 100,↓\\downarrow\)\.Boldindicates the best result;underlineindicates the second best\. Methods are grouped by training data modalities\. Values for baselines, CLOOME, InfoAlign, and CHMR are taken from\[[31](https://arxiv.org/html/2609.30433#bib.bib8),[28](https://arxiv.org/html/2609.30433#bib.bib9)\]\.MoCoPv2 achieves thebest performance on ToxCast\(70\.3±0\.270\.3\\pm 0\.2\), outperforming CHMR \(69\.3±0\.369\.3\\pm 0\.3\) by 1\.0 percentage point and InfoAlign \(66\.4±1\.166\.4\\pm 1\.1\) by 3\.9 percentage points\. Notably,MoCoPv2 achieves this using only cell morphology, whereas both CHMR and InfoAlign additionally leverage transcriptomic data during training\. On Biogen 3K,MoCoPv2 \(43\.5±0\.443\.5\\pm 0\.4MAE\) ranks second after CHMR \(40\.9±0\.340\.9\\pm 0\.3\) and substantially outperforms InfoAlign \(49\.4±0\.249\.4\\pm 0\.2\)\. On ChEMBL 2K,MoCoPv2 \(81\.5±0\.681\.5\\pm 0\.6\) matches InfoAlign and trails CHMR by 3\.2 points\. On Broad 6K,MoCoPv2 \(68\.9±1\.268\.9\\pm 1\.2\) slightly trails both InfoAlign and CHMR, but still outperforms all structure\-only baselines and CLOOME\.
Overall,MoCoPv2 outperforms InfoAlign on 3 out of 4 benchmarks \(ToxCast, ChEMBL 2K, and Biogen 3K\) and achieves the best result on ToxCast across all methods\. These results are particularly noteworthy given the architectural simplicity ofMoCoP\(a GGNN molecular encoder with contrastive alignment to DL morphological embeddings\), compared with the multi\-modal context graphs of InfoAlign or the hierarchical vector quantization of CHMR\. Where CHMR outperformsMoCoPv2 \(ChEMBL 2K, Broad 6K, and most Biogen 3K tasks\), the additional transcriptomic signal and more expressive architecture likely contribute to its advantage\. However,MoCoPv2’s competitive performance using morphology alone suggests that cell morphology captures a substantial portion of the biological information relevant to these tasks\.
Biogen 3K Per\-Task Analysis\.A per\-task breakdown of the Biogen 3K ADME regression results reveals complementary strengths \(Table[4](https://arxiv.org/html/2609.30433#Sx3.T4)\)\.MoCoPv2 achieves the best performance on 2 of 6 tasks: MDR1\-MDCK efflux ratio \(33\.8±0\.333\.8\\pm 0\.3\) and rat liver microsomal clearance \(38\.8±0\.438\.8\\pm 0\.4\), which again show the strength of the method for certain ADME predictions compared with a much more complicated architecture trained with additional transcriptomics information\. Notably,MoCoPv2 also outperforms InfoAlign on all 6 individual tasks\.
Table 4:Biogen 3K per\-task comparison \(MAE×100\\times 100,↓\\downarrow\)\. Bold indicates the best result\.
## Conclusion
We have presentedMoCoPv2, an improved molecular\-morphology contrastive pretraining framework that replaces CellProfiler\-derived morphological features with deep\-learning\-based morphological embeddings as alignment targets\. The DL embedding pipeline, comprising a fine\-tuned ResNet\-18 encoder and CORAL batch correction, extracts richer morphological representations from Cell Painting images, leading to improved molecular embeddings after alignment\.
Our key findings are: \(1\)MoCoPv2 molecular embeddings more accurately capture how molecules perturb cell morphology, even outperforming CellProfiler\-aligned embeddings on CellProfiler\-based similarity retrieval; \(2\) the pretrained embeddings consistently improve QSAR prediction in linear probe evaluations, with gains across all data proportions; \(3\)MoCoPv2 achieves state\-of\-the\-art performance on ToxCast toxicity prediction and competitive results on ADME and activity benchmarks compared to recent methods that additionally use transcriptomic data; and \(4\) retrieval performance follows a log\-linear scaling law with training data size, with consistent improvements from 10K to 98K training compounds\.
The pretrained molecular encoder fromMoCoPv2 serves as a readily tunable foundation for a broad range of*in silico*drug design applications\. As demonstrated in this work, the encoder can be efficiently adapted via either linear probing or end\-to\-end fine\-tuning to QSAR and ADMET prediction tasks that are central to lead optimization and safety assessment in drug discovery pipelines\. Because the encoder has internalized morphology\-grounded information about how molecules perturb living cells, it provides a complementary signal to purely structure\-derived descriptors, particularly for chemically novel scaffolds where structural fingerprints offer limited predictive power\. Beyond property prediction, the learned embeddings are also applicable to mechanism\-of\-action \(MOA\) retrieval: by querying nearest neighbors in theMoCoPembedding space, one can identify compounds that induce similar cellular phenotypes regardless of structural similarity, supporting scaffold hopping and target deconvolution efforts\.
Looking ahead, several avenues for further improvement become feasible\. On the DL embedding extraction side, the ResNet\-18 backbone could be replaced by self\-supervised vision transformers such as DINO\[[8](https://arxiv.org/html/2609.30433#bib.bib19),[7](https://arxiv.org/html/2609.30433#bib.bib20)\], which have demonstrated superior performance on Cell Painting data by learning attention\-based representations that capture structurally meaningful phenotypic features at both subcellular and population scales without requiring manual annotations\. Masked autoencoders\[[26](https://arxiv.org/html/2609.30433#bib.bib18),[27](https://arxiv.org/html/2609.30433#bib.bib51)\]offer another promising direction, having shown scalable improvements with larger models and datasets on microscopy images\. On the contrastive pretraining side, the current architecture treats the molecule and morphology modalities as separate encoders bridged by a cosine\-similarity objective\. An alternative approach would be to model the joint distribution of molecular tokens and morphological features through a unified transformer architecture with masked token prediction, analogous to how masked language models learn bidirectional context, which could capture richer cross\-modal dependencies than the current contrastive alignment\. The molecular encoder itself could also be upgraded from a GGNN to a more expressive architecture, such as a graph transformer\. Multi\-modal extensions incorporating transcriptomic data\[[60](https://arxiv.org/html/2609.30433#bib.bib39),[18](https://arxiv.org/html/2609.30433#bib.bib16)\]could provide additional biological context, as gene expression and morphology capture complementary aspects of cellular state\. Additionally, the current framework does not account for compound concentration: the same molecule tested at different doses can produce markedly different phenotypic responses, yetMoCoPv2 maps each compound to a single embedding regardless of dose\. Incorporating concentration\-aware embeddings, for example by conditioning the molecular encoder on the treatment dose, could enable the model to capture dose–response relationships and improve predictions for concentration\-dependent endpoints\. Similarly, cell\-line\-specific or tissue\-specific fine\-tuning of the morphological encoder could tailor the learned representations to particular biological contexts, potentially improving performance on tissue\-specific assays\. Finally, the growing availability of large\-scale Cell Painting datasets from the JUMP consortium and other sources\[[5](https://arxiv.org/html/2609.30433#bib.bib2),[13](https://arxiv.org/html/2609.30433#bib.bib48)\]offers opportunities for continued scaling, which our results suggest will yield further improvements in representation quality\.
## Acknowledgements
The authors thank Yu Yan and Lawrence Du from GSK AIML for their inspiring discussions and valuable feedbacks\. The authors also would like to thank the JUMP Cell Painting Consortium for making large\-scale Cell Painting data publicly available, and the GSK Onyx platform for support on the computation infrastructure\. This manuscript has used Claude Opus 4\.6 model for language editing and formatting assistance\.
## Data and Code Availability
The JUMP\-CP Cell Painting data are publicly available through the Cell Painting Gallery \([https://registry\.opendata\.aws/cellpainting\-gallery](https://registry.opendata.aws/cellpainting-gallery)\)\. The downstream benchmark datasets \(ChEMBL 2K, ToxCast, Broad 6K, Biogen 3K\) were obtained from the InfoAlign repository \([https://github\.com/liugangcode/InfoAlign](https://github.com/liugangcode/InfoAlign)\)\. The DL morphological embedding pipeline is a proprietary cell imaging processing workflow, and the training of MoCoP v2 reuses the codebase for MoCoP v1, which is available at[https://github\.com/GSK\-AI/mocop](https://github.com/GSK-AI/mocop)\.
## References
- \[1\]W\. Ahmad, E\. Simon, S\. Chithrananda, G\. Grand, and B\. Ramsundar\(2022\)ChemBERTa\-2: towards chemical foundation models\.arXiv preprint arXiv:2209\.01712\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2209.01712)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p1.1)\.
- \[2\]J\. Arevalo, E\. Su, J\. D\. Ewald, R\. van Dijk, A\. E\. Carpenter, and S\. Singh\(2024\)Evaluating batch correction methods for image\-based cell profiling\.Nature Communications15\(1\)\.External Links:[Document](https://dx.doi.org/10.1038/s41467-024-50613-5)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p4.1),[Batch Correction with CORAL](https://arxiv.org/html/2609.30433#Sx2.SSx2.p1.1)\.
- \[3\]I\. Bendidi, Y\. E\. Mesbahi, A\. K\. Denton, K\. Suri, K\. Kenyon\-Dean, A\. Genovesio, and E\. Noutahi\(2025\)A cross modal knowledge distillation & data augmentation recipe for improving transcriptomics representations through morphological features\.arXiv preprint arXiv:2505\.21317\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2505.21317)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p3.1)\.
- \[4\]M\. Bray, S\. M\. Gustafsdottir, M\. H\. Rohban, S\. Singh, V\. Ljosa, K\. L\. Sokolnicki, J\. A\. Bittker, N\. E\. Bodycombe, V\. Dančík, T\. P\. Hasaka, C\. S\. Hon, M\. M\. Kemp, K\. Li, D\. Walpita, M\. J\. Wawer, T\. R\. Golub, S\. L\. Schreiber, P\. A\. Clemons, A\. F\. Shamji, and A\. E\. Carpenter\(2017\)A dataset of images and morphological profiles of 30 000 small\-molecule treatments using the cell painting assay\.GigaScience6\(12\)\.External Links:[Document](https://dx.doi.org/10.1093/gigascience/giw014)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p2.1)\.
- \[5\]S\. N\. Chandrasekaran, J\. Ackerman, E\. Alix, D\. M\. Ando, J\. Arevalo, M\. Bennion, N\. Boisseau, A\. Borowa, J\. D\. Boyd, L\. Brino, P\. J\. Byrne, H\. Ceulemans, C\. Ch’ng, B\. A\. Cimini, D\. Clevert, N\. Deflaux, J\. G\. Doench, T\. Dorval, R\. Doyonnas, V\. Dragone, O\. Engkvist, P\. W\. Faloon, B\. Fritchman, F\. Fuchs, S\. Garg, T\. J\. Gilbert, D\. Glazer, D\. Gnutt, A\. Goodale, J\. Grignard, J\. Guenther, Y\. Han, Z\. Hanifehlou, S\. Hariharan, D\. Hernandez, S\. R\. Horman, G\. Hormel, M\. Huntley, I\. Icke, M\. Iida, C\. B\. Jacob, S\. Jaensch, J\. Khetan, M\. Kost\-Alimova, T\. Krawiec, D\. Kuhn, C\. Lardeau, A\. Lembke, F\. Lin, K\. D\. Little, K\. R\. Lofstrom, S\. Lotfi, D\. J\. Logan, Y\. Luo, F\. Madoux, P\. A\. M\. Zapata, B\. A\. Marion, G\. Martin, N\. J\. McCarthy, L\. Mervin, L\. Miller, H\. Mohamed, T\. Monteverde, E\. Mouchet, B\. Nicke, A\. Ogier, A\. Ong, M\. Osterland, M\. Otrocka, P\. J\. Peeters, J\. Pilling, S\. Prechtl, C\. Qian, K\. Rataj, D\. E\. Root, S\. K\. Sakata, S\. Scrace, H\. Shimizu, D\. Simon, P\. Sommer, C\. Spruiell, I\. Sumia, S\. E\. Swalley, H\. Terauchi, A\. Thibaudeau, A\. Unruh, J\. V\. de Waeter, M\. V\. Dyck, C\. van Staden, M\. Warchoł, E\. Weisbart, A\. Weiss, N\. Wiest\-Daessle, G\. Williams, S\. Yu, B\. Zapiec, M\. Żyła, S\. Singh, and A\. E\. Carpenter\(2023\)JUMP cell painting dataset: morphological impact of 136,000 chemical and genetic perturbations\.bioRxiv\.External Links:[Document](https://dx.doi.org/10.1101/2023.03.23.534023)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p2.1),[Conclusion](https://arxiv.org/html/2609.30433#Sx4.p4.1)\.
- \[6\]S\. N\. Chandrasekaran, B\. A\. Cimini, A\. Goodale, L\. Miller, M\. Kost\-Alimova, N\. Jamali, J\. G\. Doench, B\. Fritchman, A\. Skepner, M\. Melanson, A\. A\. Kalinin, J\. Arevalo, M\. Haghighi, J\. C\. Caicedo, D\. Kuhn, D\. Hernandez, J\. Berstler, H\. Shafqat\-Abbasi, D\. E\. Root, S\. E\. Swalley, S\. Garg, S\. Singh, and A\. E\. Carpenter\(2024\)Three million images and morphological profiles of cells treated with matched chemical and genetic perturbations\.Nature Methods21\(6\),pp\. 1114–1121\.External Links:[Document](https://dx.doi.org/10.1038/s41592-024-02241-6)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p2.1)\.
- \[7\]J\. O\. Cross\-Zamirski, G\. Williams, E\. Mouchet, C\. Schönlieb, R\. Turkki, and Y\. Wang\(2022\)Self\-supervised learning of phenotypic representations from cell images with weak labels\.arXiv preprint arXiv:2209\.07819\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2209.07819)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p4.1),[Conclusion](https://arxiv.org/html/2609.30433#Sx4.p4.1)\.
- \[8\]M\. Doron, T\. Moutakanni, Z\. S\. Chen, N\. Moshkov, M\. Caron, H\. Touvron, P\. Bojanowski, W\. M\. Pernice, and J\. C\. Caicedo\(2023\)Unbiased single\-cell morphology with self\-supervised vision transformers\.bioRxiv\.External Links:[Document](https://dx.doi.org/10.1101/2023.06.16.545359)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p4.1),[Conclusion](https://arxiv.org/html/2609.30433#Sx4.p4.1)\.
- \[9\]C\. Edwards, T\. Lai, K\. Ros, G\. Honke, K\. Cho, and H\. Ji\(2022\)Translation between molecules and natural language\.Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing,pp\. 375–413\.External Links:[Document](https://dx.doi.org/10.18653/v1/2022.emnlp-main.26)Cited by:[Comparing with Other State\-of\-the\-Art Methods](https://arxiv.org/html/2609.30433#Sx3.SSx4.p2.1),[Table 3](https://arxiv.org/html/2609.30433#Sx3.T3.7.4.1)\.
- \[10\]B\. Fabian, T\. Edlich, H\. Gaspar, M\. Segler, J\. Meyers, M\. Fiscato, and M\. Ahmed\(2020\)Molecular representation learning with language models and domain\-relevant auxiliary tasks\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2011.13230)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p1.1)\.
- \[11\]PyTorch lightningNote:[https://github\.com/Lightning\-AI/lightning](https://github.com/Lightning-AI/lightning)External Links:[Document](https://dx.doi.org/10.5281/zenodo.3828935)Cited by:[Molecule\-Morphology Contrastive Pretraining](https://arxiv.org/html/2609.30433#Sx2.SSx3.p6.1)\.
- \[12\]C\. Fang, Y\. Wang, R\. Grater, S\. Kapadnis, C\. Black, P\. Trapa, and S\. Sciabola\(2023\)Prospective validation of machine learning algorithms for absorption, distribution, metabolism, and excretion prediction: an industrial perspective\.Journal of Chemical Information and Modeling63\(11\),pp\. 3263–3274\.Note:PMID: 37216672External Links:[Document](https://dx.doi.org/10.1021/acs.jcim.3c00160),[Link](https://doi.org/10.1021/acs.jcim.3c00160),https://doi\.org/10\.1021/acs\.jcim\.3c00160Cited by:[Table 8](https://arxiv.org/html/2609.30433#A6.T8.5.5.1),[Comparing with Other State\-of\-the\-Art Methods](https://arxiv.org/html/2609.30433#Sx3.SSx4.p1.1)\.
- \[13\]M\. M\. Fay, O\. Kraus, M\. Victors, L\. Arumugam, K\. Vuggumudi, J\. Urbanik, K\. Hansen, S\. Celik, N\. Cernek, G\. Jagannathan, J\. Christensen, B\. A\. Earnshaw, I\. S\. Haque, and B\. Mabey\(2023\)RxRx3: phenomics map of biology\.bioRxiv\.External Links:[Document](https://dx.doi.org/10.1101/2023.02.07.527350)Cited by:[Conclusion](https://arxiv.org/html/2609.30433#Sx4.p4.1)\.
- \[14\]S\. G\. Finlayson, M\. B\.A\. McDermott, A\. V\. Pickering, S\. L\. Lipnick, and I\. S\. Kohane\(2020\)Cross\-modal representation alignment of molecular structure and perturbation\-induced transcriptional profiles\.Biocomputing 2021,pp\. 273–284\.External Links:[Document](https://dx.doi.org/10.1142/9789811232701%5F0026)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p3.1)\.
- \[15\]P\. Fradkin, P\. Azadi, K\. Suri, F\. Wenkel, A\. Bashashati, M\. Sypetkowski, and D\. Beaini\(2024\)How molecules impact cells: unlocking contrastive phenomolecular retrieval\.arXiv preprint arXiv:2409\.08302\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2409.08302)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p3.1)\.
- \[16\]A\. Gaulton, L\. J\. Bellis, A\. P\. Bento, J\. Chambers, M\. Davies, A\. Hersey, Y\. Light, S\. McGlinchey, D\. Michalovich, B\. Al\-Lazikani, and J\. P\. Overington\(2012\)ChEMBL: a large\-scale bioactivity database for drug discovery\.Nucleic Acids Research40\(D1\),pp\. D1100–D1107\.External Links:ISSN 0305\-1048,[Document](https://dx.doi.org/10.1093/nar/gkr777),[Link](https://doi.org/10.1093/nar/gkr777),https://academic\.oup\.com/nar/article\-pdf/40/D1/D1100/16955876/gkr777\.pdfCited by:[Table 8](https://arxiv.org/html/2609.30433#A6.T8.5.2.1),[Comparing with Other State\-of\-the\-Art Methods](https://arxiv.org/html/2609.30433#Sx3.SSx4.p1.1)\.
- \[17\]S\. M\. Gustafsdottir, V\. Ljosa, K\. L\. Sokolnicki, J\. A\. Wilson, D\. Walpita, M\. M\. Kemp, K\. P\. Seiler, H\. A\. Carrel, T\. R\. Golub, S\. L\. Schreiber, P\. A\. Clemons, A\. E\. Carpenter, and A\. F\. Shamji\(2013\)Multiplex cytological profiling assay to measure diverse cellular states\.PLoS ONE8\(12\),pp\. e80999\.External Links:[Document](https://dx.doi.org/10.1371/journal.pone.0080999)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p2.1)\.
- \[18\]S\. V\. Ha, S\. Jaensch, M\. M\. Kańduła, D\. Herman, P\. Czodrowski, and H\. Ceulemans\(2025\)Cross modality learning of cell painting and transcriptomics data improves mechanism of action clustering and bioactivity modelling\.Scientific Reports15\(1\)\.External Links:[Document](https://dx.doi.org/10.1038/s41598-025-05914-0)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p3.1),[Conclusion](https://arxiv.org/html/2609.30433#Sx4.p4.1)\.
- \[19\]M\. Haghighatlari, J\. Li, X\. Guan, O\. Zhang, A\. Das, C\. J\. Stein, F\. Heidar\-Zadeh, M\. Liu, M\. Head\-Gordon, L\. Bertels, H\. Hao, I\. Leven, and T\. Head\-Gordon\(2022\)NewtonNet: a newtonian message passing network for deep learning of interatomic potentials and forces\.Digital Discovery1\(3\),pp\. 333–343\.External Links:ISSN 2635\-098X,[Document](https://dx.doi.org/10.1039/d2dd00008c),[Link](https://doi.org/10.1039/d2dd00008c),https://pubs\.rsc\.org/dd/article\-pdf/1/3/333/1833540/d2dd00008c\.pdfCited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p1.1)\.
- \[20\]J\. F\. Haslum, C\. Lardeau, J\. Karlsson, R\. Turkki, K\. Leuchowius, K\. Smith, and E\. Müllers\(2024\)Cell painting\-based bioactivity prediction boosts high\-throughput screening hit\-rates and compound diversity\.Nature Communications15\(1\)\.External Links:[Document](https://dx.doi.org/10.1038/s41467-024-47171-1)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p2.1)\.
- \[21\]K\. He, X\. Zhang, S\. Ren, and J\. Sun\(2016\)Deep residual learning for image recognition\.External Links:[Document](https://dx.doi.org/10.1109/cvpr.2016.90)Cited by:[Deep\-Learning\-based Morphological Embeddings](https://arxiv.org/html/2609.30433#Sx2.SSx1.p3.1)\.
- \[22\]M\. Hofmarcher, E\. Rumetshofer, D\. Clevert, S\. Hochreiter, and G\. Klambauer\(2019\)Accurate prediction of biological assays with high\-throughput microscopy images and convolutional networks\.Journal of Chemical Information and Modeling59\(3\),pp\. 1163–1171\.External Links:[Document](https://dx.doi.org/10.1021/acs.jcim.8b00670)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p2.1)\.
- \[23\]W\. Hu, B\. Liu, J\. Gomes, M\. Zitnik, P\. Liang, V\. Pande, and J\. Leskovec\(2019\)Strategies for pre\-training graph neural networks\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.1905.12265)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p1.1)\.
- \[24\]N\. S\. Iyer, D\. J\. Michael, S\. G\. Chi, J\. Arevalo, S\. N\. Chandrasekaran, A\. E\. Carpenter, P\. Rajpurkar, and S\. Singh\(2024\)Cell morphological representations of genes enhance prediction of drug targets\.bioRxiv\.External Links:[Document](https://dx.doi.org/10.1101/2024.06.08.598076)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p2.1)\.
- \[25\]V\. Kim, N\. Adaloglou, M\. Osterland, F\. M\. Morelli, M\. Halawa, T\. König, D\. Gnutt, and P\. A\. M\. Zapata\(2025\)Self\-supervision advances morphological profiling by unlocking powerful image representations\.Scientific Reports15\(1\)\.External Links:[Document](https://dx.doi.org/10.1038/s41598-025-88825-4)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p4.1)\.
- \[26\]O\. Kraus, K\. Kenyon\-Dean, S\. Saberian, M\. Fallah, P\. McLean, J\. Leung, V\. Sharma, A\. Khan, J\. Balakrishnan, S\. Celik, D\. Beaini, M\. Sypetkowski, C\. V\. Cheng, K\. Morse, M\. Makes, B\. Mabey, and B\. Earnshaw\(2024\)Masked autoencoders for microscopy are scalable learners of cellular biology\.2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 11757–11768\.External Links:[Document](https://dx.doi.org/10.1109/CVPR52733.2024.01117)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p4.1),[Conclusion](https://arxiv.org/html/2609.30433#Sx4.p4.1)\.
- \[27\]O\. Kraus, K\. Kenyon\-Dean, S\. Saberian, M\. Fallah, P\. McLean, J\. Leung, V\. Sharma, A\. Khan, J\. Balakrishnan, S\. Celik, M\. Sypetkowski, C\. V\. Cheng, K\. Morse, M\. Makes, B\. Mabey, and B\. Earnshaw\(2023\)Masked autoencoders are scalable learners of cellular morphology\.arXiv preprint arXiv:2309\.16064\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2309.16064)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p4.1),[Conclusion](https://arxiv.org/html/2609.30433#Sx4.p4.1)\.
- \[28\]M\. Li, Z\. Zang, W\. Xing, J\. Chen, R\. Zhang, J\. Luo, and S\. Z\. Li\(2025\)Learning cell\-aware hierarchical multi\-modal representations for robust molecular modeling\.arXiv preprint arXiv:2511\.21120\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2511.21120)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p3.1),[Comparing with Other State\-of\-the\-Art Methods](https://arxiv.org/html/2609.30433#Sx3.SSx4.p2.1),[Table 3](https://arxiv.org/html/2609.30433#Sx3.T3),[Table 3](https://arxiv.org/html/2609.30433#Sx3.T3.6),[Table 3](https://arxiv.org/html/2609.30433#Sx3.T3.7.11.1),[Table 4](https://arxiv.org/html/2609.30433#Sx3.T4.5.1.3)\.
- \[29\]Y\. Li, D\. Tarlow, M\. Brockschmidt, and R\. Zemel\(2015\)Gated graph sequence neural networks\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.1511.05493)Cited by:[Molecule\-Morphology Contrastive Pretraining](https://arxiv.org/html/2609.30433#Sx2.SSx3.p2.1)\.
- \[30\]A\. Lin and A\. X\. Lu\(2022\)Incorporating knowledge of plates in batch normalization improves generalization of deep learning for microscopy images\.bioRxiv\.External Links:[Document](https://dx.doi.org/10.1101/2022.10.14.512286)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p4.1)\.
- \[31\]G\. Liu, S\. Seal, J\. Arevalo, Z\. Liang, A\. E\. Carpenter, M\. Jiang, and S\. Singh\(2024\)Learning molecular representation in a cell\.arXiv preprint arXiv:2406\.12056\.Note:ICLR 2025External Links:[Document](https://dx.doi.org/10.48550/arXiv.2406.12056)Cited by:[Appendix Appendix F](https://arxiv.org/html/2609.30433#A6.p1.1),[Introduction](https://arxiv.org/html/2609.30433#Sx1.p3.1),[Comparing with Other State\-of\-the\-Art Methods](https://arxiv.org/html/2609.30433#Sx3.SSx4.p1.1),[Comparing with Other State\-of\-the\-Art Methods](https://arxiv.org/html/2609.30433#Sx3.SSx4.p2.1),[Table 3](https://arxiv.org/html/2609.30433#Sx3.T3),[Table 3](https://arxiv.org/html/2609.30433#Sx3.T3.6),[Table 3](https://arxiv.org/html/2609.30433#Sx3.T3.7.10.1),[Table 4](https://arxiv.org/html/2609.30433#Sx3.T4.5.1.4)\.
- \[32\]S\. Liu, H\. Wang, W\. Liu, J\. Lasenby, H\. Guo, and J\. Tang\(2021\)Pre\-training molecular graph representation with 3d geometry\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2110.07728)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p1.1)\.
- \[33\]I\. Loshchilov and F\. Hutter\(2016\)SGDR: stochastic gradient descent with warm restarts\.arXiv preprint arXiv:1608\.03983\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.1608.03983)Cited by:[Molecule\-Morphology Contrastive Pretraining](https://arxiv.org/html/2609.30433#Sx2.SSx3.p6.1)\.
- \[34\]I\. Loshchilov and F\. Hutter\(2017\)Decoupled weight decay regularization\.arXiv preprint arXiv:1711\.05101\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.1711.05101)Cited by:[Deep\-Learning\-based Morphological Embeddings](https://arxiv.org/html/2609.30433#Sx2.SSx1.p4.2),[Molecule\-Morphology Contrastive Pretraining](https://arxiv.org/html/2609.30433#Sx2.SSx3.p6.1)\.
- \[35\]M\. Lu, E\. Weinberger, C\. Kim, and S\. Lee\(2025\)CellCLIP – learning perturbation effects in cell painting via text\-guided contrastive learning\.arXiv preprint arXiv:2506\.06290\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2506.06290)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p3.1)\.
- \[36\]A\. Mayr, G\. Klambauer, T\. Unterthiner, M\. Steijaert, J\. K\. Wegner, H\. Ceulemans, D\. Clevert, and S\. Hochreiter\(2018\)Large\-scale comparison of machine learning methods for drug target prediction on chembl\.Chemical Science9\(24\),pp\. 5441–5451\.External Links:[Document](https://dx.doi.org/10.1039/C8SC00148K)Cited by:[Improved QSAR Prediction on ChEMBL20](https://arxiv.org/html/2609.30433#Sx3.SSx3.p1.1)\.
- \[37\]C\. McQuin, A\. Goodman, V\. Chernyshev, L\. Kamentsky, B\. A\. Cimini, K\. W\. Karhohs, M\. Doan, L\. Ding, S\. M\. Rafelski, D\. Thirstrup, W\. Wiegraebe, S\. Singh, T\. Becker, J\. C\. Caicedo, and A\. E\. Carpenter\(2018\)CellProfiler 3\.0: next\-generation image processing for biology\.PLOS Biology16\(7\),pp\. e2005970\.External Links:[Document](https://dx.doi.org/10.1371/journal.pbio.2005970)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p4.1)\.
- \[38\]N\. Moshkov, T\. Becker, K\. Yang, P\. Horvath, V\. Dancik, B\. K\. Wagner, P\. A\. Clemons, S\. Singh, A\. E\. Carpenter, and J\. C\. Caicedo\(2023\)Predicting compound activity from phenotypic profiles and chemical structures\.Nature Communications14\(1\)\.External Links:[Document](https://dx.doi.org/10.1038/s41467-023-37570-1)Cited by:[Table 8](https://arxiv.org/html/2609.30433#A6.T8.5.4.1),[Introduction](https://arxiv.org/html/2609.30433#Sx1.p2.1),[Comparing with Other State\-of\-the\-Art Methods](https://arxiv.org/html/2609.30433#Sx3.SSx4.p1.1)\.
- \[39\]N\. Moshkov, M\. Bornholdt, S\. Benoit, M\. Smith, C\. McQuin, A\. Goodman, R\. A\. Senft, Y\. Han, M\. Babadi, P\. Horvath, B\. A\. Cimini, A\. E\. Carpenter, S\. Singh, and J\. C\. Caicedo\(2024\)Learning representations for image\-based profiling of perturbations\.Nature Communications15\(1\)\.External Links:[Document](https://dx.doi.org/10.1038/s41467-024-45999-1)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p4.1)\.
- \[40\]C\. Q\. Nguyen, D\. Pertusi, and K\. M\. Branson\(2023\)Molecule\-morphology contrastive pretraining for transferable molecular representation\.bioRxiv\.External Links:[Document](https://dx.doi.org/10.1101/2023.05.01.538999)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p5.1),[Molecule\-Morphology Contrastive Pretraining](https://arxiv.org/html/2609.30433#Sx2.SSx3.p1.1)\.
- \[41\]A\. Radford, J\. W\. Kim, C\. Hallacy, A\. Ramesh, G\. Goh, S\. Agarwal, G\. Sastry, A\. Askell, P\. Mishkin, J\. Clark, G\. Krueger, and I\. Sutskever\(2021\)Learning transferable visual models from natural language supervision\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2103.00020)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p3.1),[Molecule\-Morphology Contrastive Pretraining](https://arxiv.org/html/2609.30433#Sx2.SSx3.p4.1)\.
- \[42\]J\. Rao, H\. Lin, L\. Chen, J\. Xie, S\. Zheng, and Y\. Yang\(2025\)Multi\-modal contrastive learning with negative sampling calibration for phenotypic drug discovery\.2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 30752–30762\.External Links:[Document](https://dx.doi.org/10.1109/CVPR52734.2025.02864)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p3.1)\.
- \[43\]A\. M\. Richard, R\. S\. Judson, K\. A\. Houck, C\. M\. Grulke, P\. Volarath, I\. Thillainadarajah, C\. Yang, J\. Rathman, M\. T\. Martin, J\. F\. Wambaugh, T\. B\. Knudsen, J\. Kancherla, K\. Mansouri, G\. Patlewicz, A\. J\. Williams, S\. B\. Little, K\. M\. Crofton, and R\. S\. Thomas\(2016\)ToxCast chemical landscape: paving the road to 21st century toxicology\.Chemical Research in Toxicology29\(8\),pp\. 1225–1251\.Note:PMID: 27367298External Links:[Document](https://dx.doi.org/10.1021/acs.chemrestox.6b00135),[Link](https://doi.org/10.1021/acs.chemrestox.6b00135),https://doi\.org/10\.1021/acs\.chemrestox\.6b00135Cited by:[Table 8](https://arxiv.org/html/2609.30433#A6.T8.5.3.1),[Comparing with Other State\-of\-the\-Art Methods](https://arxiv.org/html/2609.30433#Sx3.SSx4.p1.1)\.
- \[44\]D\. Rogers and M\. Hahn\(2010\)Extended\-connectivity fingerprints\.Journal of Chemical Information and Modeling50\(5\),pp\. 742–754\.External Links:[Document](https://dx.doi.org/10.1021/ci100050t)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p1.1),[Comparing with Other State\-of\-the\-Art Methods](https://arxiv.org/html/2609.30433#Sx3.SSx4.p2.1)\.
- \[45\]Y\. Rong, Y\. Bian, T\. Xu, W\. Xie, Y\. Wei, W\. Huang, and J\. Huang\(2020\)Self\-supervised graph transformer on large\-scale molecular data\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2007.02835)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p1.1)\.
- \[46\]J\. Ross, B\. Belgodere, V\. Chenthamarakshan, I\. Padhi, Y\. Mroueh, and P\. Das\(2022\)Large\-scale chemical language representations capture molecular structure and properties\.Nature Machine Intelligence4\(12\),pp\. 1256–1264\.External Links:[Document](https://dx.doi.org/10.1038/s42256-022-00580-7)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p1.1)\.
- \[47\]A\. Sanchez\-Fernandez, E\. Rumetshofer, S\. Hochreiter, and G\. Klambauer\(2023\)CLOOME: contrastive learning unlocks bioimaging databases for queries with chemical structures\.Nature Communications14\(1\)\.External Links:[Document](https://dx.doi.org/10.1038/s41467-023-42328-w)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p3.1),[Comparing with Other State\-of\-the\-Art Methods](https://arxiv.org/html/2609.30433#Sx3.SSx4.p2.1),[Table 3](https://arxiv.org/html/2609.30433#Sx3.T3.7.7.1)\.
- \[48\]F\. Schroff, D\. Kalenichenko, and J\. Philbin\(2015\)FaceNet: a unified embedding for face recognition and clustering\.In2015 IEEE Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 815–823\.External Links:[Document](https://dx.doi.org/10.1109/cvpr.2015.7298682)Cited by:[Deep\-Learning\-based Morphological Embeddings](https://arxiv.org/html/2609.30433#Sx2.SSx1.p4.1)\.
- \[49\]K\. Schütt, O\. Unke, and M\. Gastegger\(2021\)Equivariant message passing for the prediction of tensorial properties and molecular spectra\.InProceedings of the 38th International Conference on Machine Learning,M\. Meila and T\. Zhang \(Eds\.\),Proceedings of Machine Learning Research, Vol\.139,pp\. 9377–9388\.External Links:[Link](https://proceedings.mlr.press/v139/schutt21a.html)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p1.1)\.
- \[50\]J\. Simm, G\. Klambauer, A\. Arany, M\. Steijaert, J\. K\. Wegner, E\. Gustin, V\. Chupakhin, Y\. T\. Chong, J\. Vialard, P\. Buijnsters, I\. Velter, A\. Vapirev, S\. Singh, A\. E\. Carpenter, R\. Wuyts, S\. Hochreiter, Y\. Moreau, and H\. Ceulemans\(2018\)Repurposing high\-throughput image assays enables biological activity prediction for drug discovery\.Cell Chemical Biology25\(5\),pp\. 611–618\.e3\.External Links:[Document](https://dx.doi.org/10.1016/j.chembiol.2018.01.015)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p2.1)\.
- \[51\]B\. Sun, J\. Feng, and K\. Saenko\(2016\)Return of frustratingly easy domain adaptation\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.30\.Note:ICLR 2025External Links:[Document](https://dx.doi.org/10.1609/aaai.v30i1.10306)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p5.1),[Batch Correction with CORAL](https://arxiv.org/html/2609.30433#Sx2.SSx2.p1.1)\.
- \[52\]M\. Sun, J\. Xing, H\. Wang, B\. Chen, and J\. Zhou\(2021\)MoCL: data\-driven molecular fingerprint via knowledge\-aware contrastive learning from molecular graph\.Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining,pp\. 3585–3594\.External Links:[Document](https://dx.doi.org/10.1145/3447548.3467186)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p1.1)\.
- \[53\]M\. Sypetkowski, M\. Rezanejad, S\. Saberian, O\. Kraus, J\. Urbanik, J\. Taylor, B\. Mabey, M\. Victors, J\. Yosinski, A\. R\. Sereshkeh, I\. Haque, and B\. Earnshaw\(2023\)RxRx1: a dataset for evaluating experimental batch correction methods\.2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops \(CVPRW\),pp\. 4285–4294\.External Links:[Document](https://dx.doi.org/10.1109/CVPRW59228.2023.00451)Cited by:[Batch Correction with CORAL](https://arxiv.org/html/2609.30433#Sx2.SSx2.p1.1)\.
- \[54\]G\. Tian, P\. J\. Harrison, A\. P\. Sreenivasan, J\. C\. Puigvert, and O\. Spjuth\(2022\)Combining molecular and cell painting image data for mechanism of action prediction\.bioRxiv\.External Links:[Document](https://dx.doi.org/10.1101/2022.10.04.510834)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p2.1)\.
- \[55\]M\. Trapotsi, L\. H\. Mervin, A\. M\. Afzal, N\. Sturm, O\. Engkvist, I\. P\. Barrett, and A\. Bender\(2021\)Comparison of chemical structure and cell morphology information for multitask bioactivity predictions\.Journal of Chemical Information and Modeling61\(3\),pp\. 1444–1456\.External Links:[Document](https://dx.doi.org/10.1021/acs.jcim.0c00864)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p2.1)\.
- \[56\]A\. van den Oord, Y\. Li, and O\. Vinyals\(2018\)Representation learning with contrastive predictive coding\.arXiv preprint arXiv:1807\.03748\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.1807.03748)Cited by:[Molecule\-Morphology Contrastive Pretraining](https://arxiv.org/html/2609.30433#Sx2.SSx3.p4.1)\.
- \[57\]S\. Wang, Q\. Han, W\. Qin, L\. Wang, J\. Yuan, Y\. Zhao, P\. Ren, Y\. Zhang, Y\. Tang, R\. Li, Z\. Li, W\. Zhang, S\. Gao, and F\. Bai\(2024\)PhenoScreen: a dual\-space contrastive learning framework\-based phenotypic screening method by linking chemical perturbations to cellular morphology\.bioRxiv\.External Links:[Document](https://dx.doi.org/10.1101/2024.10.23.619752)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p3.1)\.
- \[58\]Y\. Wang, J\. Wang, Z\. Cao, and A\. B\. Farimani\(2022\)Molecular contrastive learning of representations via graph neural networks\.Nature Machine Intelligence4\(3\),pp\. 279–287\.External Links:[Document](https://dx.doi.org/10.1038/s42256-022-00447-x)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p1.1)\.
- \[59\]G\. P\. Way, M\. Kost\-Alimova, T\. Shibue, W\. F\. Harrington, S\. Gill, F\. Piccioni, T\. Becker, H\. Shafqat\-Abbasi, W\. C\. Hahn, A\. E\. Carpenter, F\. Vazquez, and S\. Singh\(2021\)Predicting cell health phenotypes using image\-based morphology profiling\.Molecular Biology of the Cell32\(9\),pp\. 995–1005\.External Links:[Document](https://dx.doi.org/10.1091/mbc.E20-12-0784)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p2.1)\.
- \[60\]G\. P\. Way, T\. Natoli, A\. Adeboye, L\. Litichevskiy, A\. Yang, X\. Lu, J\. C\. Caicedo, B\. A\. Cimini, K\. Karhohs, D\. J\. Logan, M\. H\. Rohban, M\. Kost\-Alimova, K\. Hartland, M\. Bornholdt, S\. N\. Chandrasekaran, M\. Haghighi, E\. Weisbart, S\. Singh, A\. Subramanian, and A\. E\. Carpenter\(2021\)Morphology and gene expression profiling provide complementary information for mapping cell state\.bioRxiv\.External Links:[Document](https://dx.doi.org/10.1101/2021.10.21.465335)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p2.1),[Conclusion](https://arxiv.org/html/2609.30433#Sx4.p4.1)\.
- \[61\]J\. Xia, C\. Zhao, B\. Hu, Z\. Gao, C\. Tan, Y\. Liu, S\. Li, and S\. Z\. Li\(2023\)Mole\-bert: rethinking pre\-training graph neural networks for molecules\.In,External Links:[Document](https://dx.doi.org/10.26434/chemrxiv-2023-dngg4)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p1.1)\.
- \[62\]G\. Zhou, Z\. Gao, Q\. Ding, H\. Zheng, H\. Xu, Z\. Wei, L\. Zhang, and G\. Ke\(2022\)Uni\-mol: a universal 3d molecular representation learning framework\.\.External Links:[Document](https://dx.doi.org/10.26434/chemrxiv-2022-jjm0j)Cited by:[Introduction](https://arxiv.org/html/2609.30433#Sx1.p1.1),[Comparing with Other State\-of\-the\-Art Methods](https://arxiv.org/html/2609.30433#Sx3.SSx4.p2.1),[Table 3](https://arxiv.org/html/2609.30433#Sx3.T3.7.5.1)\.
## Appendix Appendix A: DL Embedding Training Hyperparameters
Table[5](https://arxiv.org/html/2609.30433#A1.T5)summarizes the key hyperparameters used for training the DL morphological embedding image encoder\.
Table 5:DL morphological embedding image encoder training hyperparameters\.#### Data Augmentation\.
Training images were augmented with random horizontal/vertical flips \(p=0\.5p=0\.5\), random 90\-degree rotations \(p=0\.5p=0\.5\), per\-channel multiplicative noise \(factors uniformly sampled from\[0\.9,1\.1\]\[0\.9,1\.1\]\), and per\-channel normalization\. Validation and test images received only per\-channel normalization\.
## Appendix Appendix B: Dataset Summary Statistics
Table[6](https://arxiv.org/html/2609.30433#A2.T6)summarizes the datasets used for DL embedding extraction, contrastive pretraining, and evaluation\. The JUMP\-CP consortium data spans 10 sources \(independent laboratories\), each contributing varying numbers of plates and compounds\. Compound\-based splitting ensures that no compound appears in both the training and test sets; the split is applied with the canonical SMILES to avoid data leakage from alternative SMILES representations\.
Table 6:Summary of datasets used in theMoCoPv2 pipeline\. “Wells” refers to the number of compound\-treated wells \(excluding DMSO controls used only for CORAL correction\)\. “Compounds” counts unique chemical entities \(by canonical SMILES\)\.StageDatasetWellsCompoundsPlatesDL image encoder trainingJUMP\-CP Source 366,47155,336217Contrastive pretraining \(DL embeddings\)Source 3 only66,47155,336217Full JUMP\-CP \(10 sources\)649,384109,1641,685Compound splits \(split 0\)Source 3 train49,785Source 3 validation2,717Source 3 test2,809Full JUMP train98,247Full JUMP validation5,459Full JUMP test5,459
## Appendix Appendix C: CORAL Batch Correction Visualization
To illustrate the necessity and effectiveness of CORAL batch correction, Figure[5](https://arxiv.org/html/2609.30433#A3.F5)shows UMAP projections of DL embeddings colored by data source, before and after CORAL correction\. Before correction, the embeddings cluster strongly by data source, indicating that batch effects dominate the representation\. After CORAL correction using DMSO negative controls, the source\-specific clusters dissolve, as can be seen in Figure[5](https://arxiv.org/html/2609.30433#A3.F5)\(b\), with data points from different sources spread out in the whole encoding space\. A small number of samples \(352 out of 47,393; 0\.7%\) with abnormally large embedding norms, caused by numerical instability in the covariance inversion for the reference source, were excluded from the CORAL\-corrected visualization\.
Figure 5:Effect of CORAL batch correction on DL morphological embeddings\. UMAP projections \(cosine distance,nneighbors=30n\_\{\\text\{neighbors\}\}=30\) colored by JUMP\-CP data source, subsampled to∼\\sim1,500 wells per source\.\(a\)Before CORAL correction: embeddings cluster by data source\. Batch effects dominate the clustering\.\(b\)After CORAL correction using DMSO negative controls: source\-specific clusters dissolve, and embeddings are intermixed across sources\. Samples with outlier embedding norms \(<<1% of data\) were excluded from the corrected panel\.
## Appendix Appendix D: MoCoP Training Hyperparameters
Table[7](https://arxiv.org/html/2609.30433#A4.T7)summarizes the key hyperparameters used forMoCoPv2 contrastive pretraining\.
Table 7:MoCoPv2 contrastive pretraining hyperparameters\.
## Appendix Appendix E: Additional MoCoP Embedding Divergence Examples
This appendix presents additional examples whereMoCoPv1 and v2 embeddings disagree on compound similarity, beyond those shown in the main text\. All compound pairs are from Source 3 \(same microscope\) to eliminate batch effects\. In all cases, visual inspection of Cell Painting images supports the v2 assessment\. Cell Painting composite: Blue = DNA, Green = AGP, and Red = Mitochondria\. Plate barcodes and well positions from JUMP\-CP are shown along with the structures\.
#### Additional strong divergence examples\.
Figure[6](https://arxiv.org/html/2609.30433#A5.F6)shows pairs where v2 correctly identifies phenotypic similarity that v1 misses\. Figure[7](https://arxiv.org/html/2609.30433#A5.F7)shows pairs where v1 incorrectly assigns high similarity to compounds with different phenotypes\.
\(a\)v2 cos=\+0\.583=\+0\.583, v1 cos=−0\.308=\-0\.308, Tanimoto=0\.135=0\.135\. Both compounds produce similar cell density and spreading patterns with concordant channel balance\.
\(b\)v2 cos=\+0\.567=\+0\.567, v1 cos=−0\.316=\-0\.316, Tanimoto=0\.170=0\.170\. Concordant phenotypes with similar cell morphology and AGP/Mito staining patterns despite low structural similarity\.
Figure 6:Additional strong divergence: v2\-similar pairs\.MoCoPv2 correctly identifies phenotypic similarity not captured by v1\.\(a\)v1 cos=\+0\.513=\+0\.513, v2 cos=−0\.459=\-0\.459, Tanimoto=0\.107=0\.107\. Compound A shows large, spread cells; Compound B shows compact, smaller cells with different morphology\.
\(b\)v1 cos=\+0\.626=\+0\.626, v2 cos=−0\.308=\-0\.308, Tanimoto=0\.126=0\.126\. Compound A shows very dense, small cells at high confluence; Compound B shows larger and more sparse cells with lower occupation of the FOV, indicating different viability\.
Figure 7:Additional strong divergence: v1\-similar pairs\.MoCoPv1 incorrectly assigns high similarity \(cos\>0\.5\>0\.5\) to compound pairs with very different phenotypes\. v2 correctly identifies these as dissimilar\.
#### Divergence examples among structurally dissimilar compounds\.
To further investigate the behavior of v1 and v2 for chemically isolated compounds, we selected pairs with low Tanimoto similarity \(<0\.2<0\.2\) and moderate embedding divergence between the two models\. Figure[8](https://arxiv.org/html/2609.30433#A5.F8)shows two types of disagreement:
- •V2\-similar pairs\(Figure[8](https://arxiv.org/html/2609.30433#A5.F8)a–b\): v2 cosine≈\+0\.80\\approx\+0\.80\(95th percentile\), v1 cosine≈\+0\.25\\approx\+0\.25\(64th percentile\)\. Visual inspection confirms concordant phenotypes \(similar cell density, spreading pattern, and channel intensity balance\) despite very different chemical structures \(Tanimoto≈0\.15\\approx 0\.15\)\.MoCoPv2 correctly identifies this morphological similarity, demonstrating its utility for finding phenotypically related compounds across distinct scaffolds\.
- •V1\-similar pairs\(Figure[8](https://arxiv.org/html/2609.30433#A5.F8)c–d\): v1 cosine≈\+0\.83\\approx\+0\.83\(96th percentile\), v2 cosine≈\+0\.18\\approx\+0\.18\(61st percentile\)\. Visual inspection reveals visibly different phenotypes with clear differences in cell density, size, morphology, and staining patterns\.MoCoPv1 incorrectly assigns high similarity; v2’s lower score is more appropriate\.
\(a\)V2\-similar, example 1 \(Tanimoto=0\.152=0\.152, v2 cos=\+0\.789=\+0\.789, v1 cos=\+0\.254=\+0\.254\)\. Both compounds show similar cell density and spreading with concordant channel proportions\.
\(b\)V2\-similar, example 2 \(Tanimoto=0\.172=0\.172, v2 cos=\+0\.800=\+0\.800, v1 cos=\+0\.257=\+0\.257\)\. Concordant morphology with similar cell size and distribution patterns\.
\(c\)V1\-similar, example 1 \(Tanimoto=0\.161=0\.161, v1 cos=\+0\.835=\+0\.835, v2 cos=\+0\.184=\+0\.184\)\. Compound A shows sparse, large cells with irregular morphology; Compound B shows dense, uniformly distributed smaller cells\. v1 incorrectly rates these as similar\.
\(d\)V1\-similar, example 2 \(Tanimoto=0\.160=0\.160, v1 cos=\+0\.812=\+0\.812, v2 cos=\+0\.139=\+0\.139\)\. Compound A shows sparse cells with bright mitochondrial staining; Compound B shows a denser field with different morphology and staining balance\.
Figure 8:Divergence examples among structurally dissimilar compounds \(Tanimoto<0\.2<0\.2\), imaged on the same microscope \(Source 3\)\.\(a–b\)V2\-similar: v2 correctly identifies phenotypic similarity that v1 misses, demonstrating the value of DL\-embedding\-aligned representations for identifying shared biology across distinct scaffolds\.\(c–d\)V1\-similar: v1 incorrectly assigns high similarity to pairs with visibly different phenotypes; v2’s lower score is more appropriate\.
## Appendix Appendix F: Downstream Task Training Details
For all downstream benchmark tasks, we followed the evaluation protocol established by InfoAlign\[[31](https://arxiv.org/html/2609.30433#bib.bib8)\]\. Datasets were obtained from the InfoAlign repository, with scaffold\-based splitting used for train/validation/test partitioning\. Classification tasks used AUROC as the primary metric; regression tasks \(Biogen 3K\) used MAE\. All results are reported as 3\-seed mean±\\pmstandard deviation on the test set, by training separate models initialized from the pretrained checkpoint\.
Table[8](https://arxiv.org/html/2609.30433#A6.T8)summarizes the benchmark datasets\.
Table 8:Summary of downstream benchmark datasets\.#### Hyperparameter Sweep\.
A comprehensive hyperparameter sweep was conducted across learning rate \(5×10−55\\times 10^\{\-5\}to10−310^\{\-3\}\), dropout \(0\.1–0\.3\), early stopping patience \(10–20 epochs\), optimizer \(Adam vs\. AdamW\), and learning rate scheduler \(constant vs\. cosine annealing\)\. Key findings from the sweep:
- •Higher learning rates \(3×10−4\\times 10^\{\-4\}to 7×10−4\\times 10^\{\-4\}\) consistently improved validation performance over the baseline \(5×10−55\\times 10^\{\-5\}\)\.
- •Reduced early stopping patience \(10 vs\. 20 epochs\) reduced overfitting to the validation set\.
- •Moderate dropout \(0\.2\) generalized better than 0\.1 at high learning rates; 0\.3 hurt validation performance\.
- •Cosine annealing provided marginal improvement for ChEMBL 2K only \(\+0\.3% AUC\)\.
- •AdamW weight decay did not outperform Adam\.
Table[9](https://arxiv.org/html/2609.30433#A6.T9)reports the best configuration for each dataset\.
Table 9:Best hyperparameter configurations for each downstream dataset\.相似文章
用于多任务ADME性质预测的概率对比预训练
本文提出了一种用于分子图变换器的概率对比预训练框架,以改善药物发现中的多任务ADME性质预测,在三个基准上取得了显著提升。
基于结构保持的细胞表型分子表示学习
本文提出了 PhenMol,一个保持结构的表型感知分子表示学习框架,它在保留化学结构邻域组织的同时整合细胞表型信息,从而改进分子性质预测和药物发现任务。
用于分子属性预测的程序化预训练
本文介绍了一种三阶段训练流水线,通过程序化预训练提升分子属性预测性能,在数据稀缺情况下通过从抽象生成数据中学习归纳偏置,表现出更优性能。
MolEmb:多模态大型语言模型可以成为强大的分子嵌入模型
本文介绍了MolEmb,这是一个轻量级框架,它调整多模态大型语言模型以实现通用的分子嵌入,支持上下文感知的表示和跨模态检索,同时还提出了一个诊断基准MolCAR。
基于图神经网络、深度交叉网络和SMILES嵌入的多模态分子表示学习
本文提出了一种三分支模块化融合神经网络,整合了3D几何结构、SMILES嵌入和物理化学描述符用于分子性质预测,在QM9数据集上以不到一百万个参数实现了20.6%的误差降低。