EnSol:一种用于分子溶解度预测的环境感知图神经网络

arXiv cs.LG 论文

摘要

EnSol是一种环境感知图神经网络,它通过将溶质和溶剂表示为图,并使用交叉注意力来建模相互作用,同时利用概率输出来捕捉温度效应和实验不确定性,从而预测分子溶解度。该模型在基准数据集上取得了最先进的性能,并经过实验验证。

arXiv:2609.21151v1 Announce Type: new Abstract: Molecular solubility directly affects key aspects of molecular development such as reaction feasibility, formulation performance, separation efficiency, and solvent selection. However, experimental measurement across solutes, solvents, and temperatures remains costly and sparsely sampled. Existing computational models often rely on fixed-solvent assumptions, deterministic formulations, or simplified representations of solute-solvent interactions, limiting their ability to capture complex molecular interactions, continuous temperature effects, and experimental uncertainty. Here, we introduce EnSol, an environment-aware probabilistic framework for molecular solubility prediction. EnSol represents the solute and solvent as molecular graphs and learns separate representations for each before bringing them together through cross-attention to capture solute-solvent interactions. Temperature is incorporated directly into the solvent environment through feature-wise modulation, and a mixture density network predicts full solubility distributions to capture both temperature-dependent behavior and experimental uncertainty. On the independent SolProp and Leeds benchmark datasets, EnSol achieved Spearman correlations of 0.876 and 0.601, respectively, outperforming state-of-the-art solubility prediction models across both benchmarks. Beyond computational benchmarking, experimental validation across chemically diverse solute-solvent pairs showed that EnSol maintained strong predictive performance and supported reliable solvent ranking, achieving a Spearman correlation of 0.715. These results show that EnSol can support reliable solubility prediction and solvent selection across diverse chemical systems while accounting for predictive uncertainty.
查看原文
查看缓存全文

缓存时间: 2026/09/21 09:24

# EnSol: an environment-aware graph neural network for molecular solubility prediction
Source: [https://arxiv.org/html/2609.21151](https://arxiv.org/html/2609.21151)
\\miscnote

\*To whom correspondence should be addressed\. Tel: \(217\) 244\-0862, Email: hengji@illinois\.edu \(H\.J\.\); \(217\) 333\-2631, Email: zhao5@illinois\.edu \(H\.Z\.\)\.

Thao NguyenAffiliation:Siebel School of Computing and Data Science, University of Illinois Urbana\-Champaign, Urbana, IL, 61801, USAAffiliation:Carl R\. Woese Institute for Genomic Biology, University of Illinois Urbana\-Champaign, Urbana, IL, 61801, USAAffiliation:NSF Molecule Maker Lab Institute, University of Illinois Urbana\-Champaign, Urbana, IL, 61801, USASaman ShafaeiAffiliation:Carl R\. Woese Institute for Genomic Biology, University of Illinois Urbana\-Champaign, Urbana, IL, 61801, USAAffiliation:Department of Chemical and Biomolecular Engineering, University of Illinois Urbana\-Champaign, Urbana, IL, 61801, USAAffiliation:DOE Center for Advanced Bioenergy and Bioproducts Innovation, University of Illinois Urbana\-Champaign, Urbana, IL, 61801, USAZhengyi ZhangAffiliation:Carl R\. Woese Institute for Genomic Biology, University of Illinois Urbana\-Champaign, Urbana, IL, 61801, USAAffiliation:Department of Chemical and Biomolecular Engineering, University of Illinois Urbana\-Champaign, Urbana, IL, 61801, USAAffiliation:DOE Center for Advanced Bioenergy and Bioproducts Innovation, University of Illinois Urbana\-Champaign, Urbana, IL, 61801, USAHuimin ZhaoAffiliation:Carl R\. Woese Institute for Genomic Biology, University of Illinois Urbana\-Champaign, Urbana, IL, 61801, USAAffiliation:Department of Chemical and Biomolecular Engineering, University of Illinois Urbana\-Champaign, Urbana, IL, 61801, USAAffiliation:NSF Molecule Maker Lab Institute, University of Illinois Urbana\-Champaign, Urbana, IL, 61801, USAAffiliation:DOE Center for Advanced Bioenergy and Bioproducts Innovation, University of Illinois Urbana\-Champaign, Urbana, IL, 61801, USAHeng JiAffiliation:Siebel School of Computing and Data Science, University of Illinois Urbana\-Champaign, Urbana, IL, 61801, USAAffiliation:Carl R\. Woese Institute for Genomic Biology, University of Illinois Urbana\-Champaign, Urbana, IL, 61801, USAAffiliation:NSF Molecule Maker Lab Institute, University of Illinois Urbana\-Champaign, Urbana, IL, 61801, USAAffiliation:DOE Center for Advanced Bioenergy and Bioproducts Innovation, University of Illinois Urbana\-Champaign, Urbana, IL, 61801, USA

###### Abstract

Molecular solubility directly affects key aspects of molecular development such as reaction feasibility, formulation performance, separation efficiency, and solvent selection\. However, experimental measurement across solutes, solvents, and temperatures remains costly and sparsely sampled\. Existing computational models often rely on fixed\-solvent assumptions, deterministic formulations, or simplified representations of solute–solvent interactions, limiting their ability to capture complex molecular interactions, continuous temperature effects, and experimental uncertainty\. Here, we introduce EnSol, an environment\-aware probabilistic framework for molecular solubility prediction\. EnSol represents the solute and solvent as molecular graphs and learns separate representations for each before bringing them together through cross\-attention to capture solute–solvent interactions\. Temperature is incorporated directly into the solvent environment through feature\-wise modulation, and a mixture density network predicts full solubility distributions to capture both temperature\-dependent behavior and experimental uncertainty\. On the independent SolProp and Leeds benchmark datasets, EnSol achieved Spearman correlations of 0\.876 and 0\.601, respectively, outperforming state\-of\-the\-art solubility prediction models across both benchmarks\. Beyond computational benchmarking, experimental validation across chemically diverse solute–solvent pairs showed that EnSol maintained strong predictive performance and supported reliable solvent ranking, achieving a Spearman correlation of 0\.715\. These results show that EnSol can support reliable solubility prediction and solvent selection across diverse chemical systems while accounting for predictive uncertainty\.

###### keywords

deep learning, mixture density network, solubility prediction, solvent screening

††equal\-contributors:These authors contributed equally to this work\.††equal\-contributors:These authors contributed equally to this work\.††equal\-contributors:These authors contributed equally to this work\.## 1Main

Molecular solubility is an important physicochemical property that influences the behavior and utility of small molecules across different solvents and environments\.\[[1](https://arxiv.org/html/2609.21151#bib.bib1),[2](https://arxiv.org/html/2609.21151#bib.bib2),[3](https://arxiv.org/html/2609.21151#bib.bib3)\]In practice, insufficient or poorly characterized molecular solubility constrains formulation stability, mass transfer efficiency, and reaction performance, often increasing experimental iteration and development timelines\.\[[4](https://arxiv.org/html/2609.21151#bib.bib4),[5](https://arxiv.org/html/2609.21151#bib.bib5),[6](https://arxiv.org/html/2609.21151#bib.bib6)\]In industrial and laboratory settings, suboptimal solubility can necessitate additional burden such as phase management, and downstream processing steps, increasing both operational complexity and cost\.\[[7](https://arxiv.org/html/2609.21151#bib.bib7),[8](https://arxiv.org/html/2609.21151#bib.bib8)\]Despite its key role, systematic experimental characterization of molecular solubility across diverse solvent systems and temperature conditions remains labor\-intensive, time\-consuming, and costly, making large\-scale screening difficult during early molecular development and process design\.\[[9](https://arxiv.org/html/2609.21151#bib.bib9),[10](https://arxiv.org/html/2609.21151#bib.bib10),[11](https://arxiv.org/html/2609.21151#bib.bib11)\]Consequently, reliable predictive models for molecular solubility offer a practical way to accelerate decision\-making, reduce experimental bottlenecks, and support rational solvent and process selection\.\[[12](https://arxiv.org/html/2609.21151#bib.bib12),[13](https://arxiv.org/html/2609.21151#bib.bib13)\]

Recently, machine learning \(ML\) has been increasingly applied to predict molecular solubility and prioritize solvents before extensive laboratory testing\.\[[14](https://arxiv.org/html/2609.21151#bib.bib14)\],\[[15](https://arxiv.org/html/2609.21151#bib.bib15)\],\[[16](https://arxiv.org/html/2609.21151#bib.bib16)\]However, existing models often oversimplify this problem in ways that limit their ability for realistic solvent selection\. Many approaches focus on aqueous solubility or fixed\-solvent settings, whereas practical chemical workflows require comparison across many solvent environments\.\[[3](https://arxiv.org/html/2609.21151#bib.bib3)\],\[[17](https://arxiv.org/html/2609.21151#bib.bib17)\]Other models incorporate solvent information but treat solute and solvent representations as static features, limiting their ability to learn specific interaction patterns among solvent and solutes\.\[[18](https://arxiv.org/html/2609.21151#bib.bib18)\],\[[19](https://arxiv.org/html/2609.21151#bib.bib19)\]Furthermore, despite the direct influence of temperature on molecular dissolution and solvent behavior, it is often treated as a simple input variable rather than as a continuous environmental factor that modulates the solvation environment and solute–solvent interactions\. Lastly, most solubility predictors return a single deterministic point as prediction, even though experimental measurements can vary across protocols, datasets, and thermodynamic conditions\. As a result, these models provide little information about prediction uncertainty, making it harder to judge which predictions are reliable enough to guide experiments\.\[[6](https://arxiv.org/html/2609.21151#bib.bib6)\],\[[20](https://arxiv.org/html/2609.21151#bib.bib20)\]

To address these limitations, we developed EnSol, an environment\-aware probabilistic ML\-based model for molecular solubility prediction\. EnSol is designed around the idea that accurate solubility prediction requires modeling the full experimental context, including the solute, solvent, and temperature\. The model represents solute and solvent molecules as molecular graphs and encodes them using graph neural network \(GNN\)\. Their learned representations are then coupled through cross\-attention, allowing the model to construct interaction\-aware features that depend on the specific solute–solvent pair\. Temperature is introduced through modulation of the solvent representation, so changes in temperature can directly influence how the solvent environment is represented by the model\. EnSol also moves beyond a single\-point prediction to model a full distribution of possible solubility values, allowing the model to both account for heterogeneous or multimodal behavior and provide uncertainty estimates that help distinguish more confident predictions from those that may require additional experimental testing\.

EnSol was designed to reflect the experimental setting as closely as possible by explicitly modeling the solubility task\. We evaluate EnSol through two separate computational benchmarks and an independent experimental solvent\-screening validation\. Across all settings, EnSol consistently outperformed the comparison models, supporting its use for practical solubility prediction and solvent selection, showing how the model can move beyond retrospective prediction toward practical use in solvent screening by directly incorporating molecular context and environmental conditions into data\-driven molecular solubility prediction\.

## 2Results

### 2\.1Overview of the EnSol architecture

EnSol is an environment\-aware molecular solubility prediction framework that jointly models solute structure, solvent structure, and temperature within a unified deep learning architecture \(Fig\.[1](https://arxiv.org/html/2609.21151#S2.F1)\)\. Solute and solvent molecules are represented as molecular graphs and encoded using GNN modules, after which their node\-level representations are coupled through cross\-attention to learn context\-dependent solute–solvent interaction features\. Temperature is incorporated as a continuous environmental variable that modulates the solvent embedding through feature\-wise linear modulation \(FiLM\)\[[21](https://arxiv.org/html/2609.21151#bib.bib21)\], producing a unified solute–solvent–temperature representation of the experimental condition\. This joint representation parameterizes a Mixture Density Network\(MDN\)\[[22](https://arxiv.org/html/2609.21151#bib.bib22)\], letting EnSol to predict full conditional solubility distributions rather than a point estimates\. With integrating molecular interaction, continuous temperature conditioning, and probabilistic inference, EnSol captures heterogeneous solubility behavior across solvent and temperature regimes and provides uncertainty estimates for downstream screening and decision\-making\.

![Refer to caption](https://arxiv.org/html/2609.21151v1/figures/figure1.png)Figure 1:a\)Conceptual illustration of the solubility landscape, showing how a single solute exhibits condition\-dependent solubility across different solvents and temperatures, motivating environment\-aware modeling\.b\)Overview of the EnSol architecture\. Solute and solvent molecules are encoded as graphs, coupled via cross\-attention, and conditioned on temperature to form an environment embedding\. A mixture density network predicts a full solubility distribution, enabling uncertainty\-aware predictions\.
### 2\.2Training and comparative evaluation of EnSol

EnSol was trained on BigSolDB\[[23](https://arxiv.org/html/2609.21151#bib.bib23)\], one of the largest and most complete databases of experimentally measured molecular solubility values across diverse organic solvents and temperatures \(Fig\. S1\)\. Solute\-level splitting was used for training and validation to avoid leakage, and performance was evaluated on two independent external benchmark datasets, SolProp\[[24](https://arxiv.org/html/2609.21151#bib.bib24)\], and Leeds\[[6](https://arxiv.org/html/2609.21151#bib.bib6)\], after removing solutes overlapping with the training set \(Fig\. S2 and Fig\. S3\)\.

We benchmarked EnSol against two state\-of\-the\-art solubility prediction models, FASTSOLV\[[19](https://arxiv.org/html/2609.21151#bib.bib19)\]and Vermeire\[[24](https://arxiv.org/html/2609.21151#bib.bib24)\], using identical filtered test sets\. EnSol performance is across three random seeds, whereas both baseline models are pretrained and deterministic\. On the SolProp benchmark, EnSol showed a substantial improvement in solubility ranking, achieving a Spearman correlation of 0\.876±\\pm0\.005 compared with 0\.509 for FASTSOLV and 0\.569 for the Vermeire model \(Fig\.[2](https://arxiv.org/html/2609.21151#S2.F2)a\), while reducing RMSE to 0\.824±\\pm0\.032 from 1\.303 and 1\.700, respectively \(Fig\.[2](https://arxiv.org/html/2609.21151#S2.F2)b, TableS1 and Fig\. S4a\)\.

On the Leeds benchmark, which is more distant from the training distribution, EnSol clearly outperformed the Vermeire model while closely matching FASTSOLV\. EnSol achieved a Spearman correlation of 0\.602±\\pm0\.013, compared with 0\.593 for FASTSOLV and 0\.245 for the Vermeire model \(Fig\.[2](https://arxiv.org/html/2609.21151#S2.F2)a\)\. A similar pattern was observed for prediction error, as EnSol achieved a Root Mean Square Error \(RMSE\) of 0\.944±\\pm0\.017, which was comparable to 0\.922 for FASTSOLV and substantially lower than 2\.026 for the Vermeire model \(Fig\.[2](https://arxiv.org/html/2609.21151#S2.F2)b, TableS1and Fig\. S4b\)\.

![Refer to caption](https://arxiv.org/html/2609.21151v1/figures/figure2.png)Figure 2:Predictive accuracy and ranking performance across two independent benchmark datasets SolProp and Leeds\. Model performance is compared on two external test sets, SolProp and Leeds\. Values above bars are mean values\. EnSol error bars are standard deviation across three random seeds \(n = 3\) and FASTSOLV and the Vermeire model are pretrained and deterministic\.a\)Spearman rank correlationb\)RMSE values\.Because practical solvent selection often requires ranking candidate solvents for a fixed solute under defined experimental conditions, we further evaluated EnSol in a solvent\-ranking setting using the SolProp database\.\[[25](https://arxiv.org/html/2609.21151#bib.bib25)\]Candidate solvents were prioritized for each solute–temperature pair according to the predicted solubility\. Compared to FASTSOLV\[[19](https://arxiv.org/html/2609.21151#bib.bib19)\], EnSol identified high\-solubility solvents among the top\-ranked candidates more consistently and better maintained the relative ordering of solvent performance, reaching top 1 recall of 0\.61 and top 3 recall of 0\.85 versus 0\.35 and 0\.56 respectively \(Fig\. S5\)\. This suggests EnSol can support solvent screening as a decision\-making tool, rather than only serving as a point predictor of individual solubility measurements\.

### 2\.3Transferability and expert model fine\-tuning

Solubility models are often adapted to settings in which the relevant chemical space of solutes or solvent are narrow and only a limited number of measurements are available, such as a new compound series or a single solvent of interest\. We therefore examined whether the representations learned by EnSol from the chemically diverse, multi\-solvent BigSolDB dataset could transfer to aqueous solubility prediction, and whether the value of this pretraining correlates with the number and composition of the target data\. To this objective, transfer learning capability of EnSol was evaluated on two complementary aqueous datasets\. We first incorporated AqSolDB\[[26](https://arxiv.org/html/2609.21151#bib.bib26)\], a large curated collection of experimentally measured aqueous solubilities\. After removing compounds that overlapped with the aqueous subset of BigSolDB, 9,748 unique compounds remained\. We also included ESOL\[[12](https://arxiv.org/html/2609.21151#bib.bib12)\], a widely used aqueous solubility dataset containing 1,082 compounds\. For each dataset, we trained EnSol in two ways\. For each dataset, we compared fine\-tuning EnSol from a BigSolDB\-pretrained checkpoint with training the same architecture from random initialization using identical Murcko\-scaffold splits and optimization settings across five seeds\. Because the target datasets do not report temperature, both approaches used a temperature\-independent version of EnSol\. We additionally evaluated fixed subsets of 1,000 AqSolDB compounds to examine how transfer changed with target data availability\.

Across both datasets, the benefit of pretraining was more apparent in RMSE than in Spearman correlation, suggesting that pretraining mainly helped EnSol predict solubility values more accurately, while having a more modest effect on the relative ranking of compounds\. On ESOL, pretraining increased the Spearman correlation from 0\.922±\\pm0\.029 to 0\.942±\\pm0\.016 \(Fig\.[3](https://arxiv.org/html/2609.21151#S2.F3)a\) while it had larger effect on the RMSE and reduced the RMSE from 0\.874±\\pm0\.066 to 0\.759±\\pm0\.065 \(Fig\.[3](https://arxiv.org/html/2609.21151#S2.F3)b\)\. The benefit was also followed the same pattern in the data\-limited AqSolDB setting\. With 1,000 fine\-tuning compounds, pretraining reduced the RMSE from 1\.383±\\pm0\.181 to 1\.292±\\pm0\.139 and increased the Spearman correlation from 0\.783±\\pm0\.039 to 0\.804±\\pm0\.038\. Also, as more target\-domain data were introduced, the contribution of pretraining gradually shrank\. On the full AqSolDB dataset, pretraining produced only a small reduction in RMSE, from 1\.146±\\pm0\.070 to 1\.102±\\pm0\.143, while the Spearman correlation changed little, from 0\.864±\\pm0\.028 to 0\.869±\\pm0\.032 \(Table S2\)\. None of these differences reached statistically significant, indicating that the observed gains were modest relative to variation arising from initialization and scaffold partitioning and relationship between target dataset size and transfer benefit was not strictly monotonic\.

The contrast between ESOL and AqSolDB further indicates that transfer depends on the composition and internal consistency of the target dataset rather than on sample size alone\. Despite containing approximately nine\-fold fewer compounds, ESOL supported substantially lower prediction error and higher rank correlation than the full AqSolDB dataset\. This difference is consistent with the narrower and more uniformly curated chemical space of ESOL, whereas AqSolDB aggregates measurements from multiple sources and covers a broader and more heterogeneous distribution\. Overall, the results indicate that BigSolDB pretraining provides a slight transferable initialization for aqueous solubility prediction, with the clearest gains observed when target measurements are limited\. The benefit is more consistent for reducing absolute prediction error than for improving compound ranking and becomes less pronounced as sufficient task\-specific data become available\.

![Refer to caption](https://arxiv.org/html/2609.21151v1/figures/figure3.png)Figure 3:Transfer learning and generalization of EnSol\. Fine\-tuning from a BigSolDB\-pretrained checkpoint compared with training the identical architecture from random initialization, across three random seeds \(n = 3\) of the two deduplicated AqSolDB set \(n = 1,000 and n= 9,748\) probe the effect of target dataset size and ESOL \(n = 1,082\) is an independent aqueous dataset shown for comparison\. Values above bars are means\. a\) Spearman rank correlation b\) RMSE values\.
### 2\.4Ablation studies and model component contributions

To quantify the contribution of each individual architectural component within EnSol, we performed a systematic ablation study on SolProp test set, in which key modules of the model were selectively removed or replaced while keeping the rest of the parameters including training splits, optimization settings, and evaluation protocol fixed over three different seeds\. This comparison allowed us to isolate the functional contribution of molecular encoding, solute–solvent interaction modeling via cross\-attention, temperature conditioning, and probabilistic prediction to overall model performance\.

The results indicated that molecular encoder and solute–solvent interaction mechanism made large contributions to performance\. Replacing AttentiveFP\[[26](https://arxiv.org/html/2609.21151#bib.bib26)\]with a conventional MPNN\[[27](https://arxiv.org/html/2609.21151#bib.bib27)\], corresponded to an average decrease of 0\.088 in Spearman correlation and an increase of 0\.187 in RMSE, which showcases the importance of the way the solvents and solute molecules are represented\. Removing solute–solvent cross\-attention also decreased Spearman correlation by 0\.088 and increased RMSE by 0\.104, showing that explicitly modeling interactions between the two molecular representations improves performance\. Replacing FiLM\-based temperature conditioning with direct temperature concatenation decreased Spearman correlation by 0\.042 and increased RMSE by 0\.092, further supporting our hypothesis of temperature\-dependent modulation of the solvent representation\. Similarly, replacing the mixture\-density network with a deterministic mean squared error \(MSE\) regression head decreased Spearman correlation by 0\.043 and increased RMSE by 0\.100 \(Fig\.[4](https://arxiv.org/html/2609.21151#S2.F4)a, Fig\.[4](https://arxiv.org/html/2609.21151#S2.F4)b and Table S3\)\. The combination of the ablation study results suggest that each major component contributes to EnSol performance, with the molecular encoder and solute–solvent cross\-attention producing the largest improvements in ranking performance\.

![Refer to caption](https://arxiv.org/html/2609.21151v1/figures/figure4.png)Figure 4:Component\-wise ablation of the EnSol architecture on the SolProp benchmark\. Each variant modifies exactly one design choice relative to full EnSol, holding training splits, optimization settings, and evaluation protocol fixed\. K = 5 components, the Gaussian mixture is increased from three to five components; No cross\-attention, the solute–solvent cross\-attention layer is removed; MSE head, the mixture\-density head is replaced by a deterministic regression head trained with mean\-squared error; Concat\. temperature, FiLM temperature conditioning is replaced by a normalized temperature scalar concatenated onto the solute–solvent embedding; MPNN encoder, the AttentiveFP graph encoder is replaced by a message\-passing neural network\. Performance is shown as a\) Spearman rank correlation\. b\) RMSE values\.
### 2\.5Experimental validation and solvent recommendation

Although computational benchmarks establish comparative predictive performance, molecular solubility remains an experimentally measured property, and practical utility requires that model predictions translate into reliable solvent selection under fixed laboratory conditions\. To evaluate this capability, we performed targeted experimental validation, testing whether EnSol could prioritize suitable solvents beyond in silico benchmarking across a diverse set of solutes \(Fig\.[5](https://arxiv.org/html/2609.21151#S2.F5)a\)\. We experimentally measured solubility across 100 chemically diverse solute–solvent pairs, comprising 10 solutes and 10 commonly used solvents selected to reflect realistic solvent\-selection scenarios \(Fig\. S6, Table S4\)\. Out of these, 78 yielded quantitative solubility values while the remaining 22 fell below the 1 mg mL\-1detection limit of the visual assay and are therefore excluded from the quantitative metrics reported while being retained in the deposited dataset \(Fig\.[5](https://arxiv.org/html/2609.21151#S2.F5)b\)\.

Compared with the computational benchmarks, this experimental dataset provides a more controlled test of model performance, as all solubility measurements were generated under a consistent experimental protocol and within a fixed experimental setting\. Under these conditions, EnSol showed stronger agreement with measured solubilities than the comparison models \(Table S5, Fig\. S7\)\. EnSol achieved a Spearman correlation of 0\.715±\\pm0\.035, compared with 0\.31 for FASTSOLV and 0\.595 for the Vermeire model \(Fig\.[5](https://arxiv.org/html/2609.21151#S2.F5)c\)\. The same trend was observed for prediction error, with EnSol reaching an RMSE of 0\.699±\\pm0\.062, compared with 1\.038 for FASTSOLV and 0\.841 for the Vermeire model \(Fig\.[5](https://arxiv.org/html/2609.21151#S2.F5)d\)\. Furthermore, EnSol more reliably prioritized experimentally favorable solvents compared with the baseline model\. For solutes measured across multiple solvents, EnSol recovered the experimental solvent ranking with high fidelity for the majority of compounds, achieving per\-solute Spearman correlations above 0\.7 for six of the ten solutes and reaching 1\.00 for threonine, 0\.98 for thiourea, and 0\.88 for 8\-hydroxyquinoline \(Fig\. S8\)\. However, this performance was not uniform across all compounds, and some solutes remained challenging for the model\. Performance was weaker for a small number of solutes, most notably the phosphorane \(Wittig ylide\) and chlorothioxanthone\. This is likely because both compounds are structurally unusual relative to typical solubility training data, with the phosphorane in particular representing a reactive organophosphorus chemistry that is poorly represented in existing datasets, leading to less reliable extrapolation\. Such cases highlight an important limitation of the current model and the need for broader prospective validation across more diverse chemistries and experimental conditions\.

![Refer to caption](https://arxiv.org/html/2609.21151v1/figures/figure5.png)Figure 5:Experimental validation set design and model performance\. a\)Structural diversity of the 10 solutes used for experimental validation, spanning a range of functional groups, sizes, and scaffolds\.b\)The validation set spans 10 solutes and 10 solvents, all evaluated at 298\.15 K, filled circles mark the solute–solvent pairs included in the set and crossed\-out circles mark pairs that were not due to the measurements below the 1 mg mL\-1detection limit of the visual assay\.c\)Model performance is compared on the experimental validation set\. Values above bars are mean values\. EnSol error bars are standard deviation across three random seeds \(n = 3\) and FASTSOLV and the Vermeire model are pretrained and deterministic across Spearman rank correlationd\)Model performance is compared on the experimental validation set\. Values above bars are mean values\. EnSol error bars are standard deviation across three random seeds \(n = 3\) and FASTSOLV and the Vermeire model are pretrained and deterministic across RMSE values\.

## 3Conclusions

Solubility is one of the fundamental properties of molecules which directly affects how compounds behave under specific solvent and temperature conditions\. In this work, we developed EnSol, an environment\-aware probabilistic framework for molecular solubility prediction, by jointly modeling solute structure, solvent structure, and temperature\-dependent environmental context\. Through systematic computational evaluation on two independent external benchmarks, we showed that EnSol improves solubility ranking on the SolProp benchmark and matches the strongest available baseline under greater distribution shift\. Furthermore, experimental validation across chemically diverse solutes and candidate solvents confirmed that EnSol predictions translate into practical solvent\-ranking capability, to reduce experimental screening burden and improve the efficiency of early\-stage molecular and process development\.

We anticipate that EnSol will serve as a useful platform for data\-driven solubility prediction tool in pharmaceutical development, synthetic chemistry, formulation design, and sustainable process engineering\. By predicting solubility distributions, EnSol enables users to prioritize solvent candidates for experimental testing, thereby reducing screening burden, material consumption, and iteration time\. Moreover, the transferability of EnSol representations provides a practical route for building expert solubility models when additional scars private or domain\-specific measurements are available for a particular solute class, solvent family, or application regime\.

However, despite EnSol’s capabilities, several limitations remain, including limited mechanistic interpretability, dependence on the quality and coverage of available solubility measurements, and reduced confidence in sparsely sampled regimes such as uncommon solutes, extreme temperatures, and solvent mixtures\. Looking ahead, integrating EnSol with automated experimentation, active learning, and closed\-loop solvent optimization could enable continuously improving solubility models that directly connect predictive uncertainty, experimental feasibility, and sustainability\-aware chemical design\.

## 4Methods

### 4\.1Train and test databases

EnSol was trained on BigSolDB\[[23](https://arxiv.org/html/2609.21151#bib.bib23)\], a well\-curated database of experimentally measured solubility values for organic compounds across a broad range of solvents and temperatures\. Solubility values were standardized as log values of the solubility measurements\. The final training dataset contained 100,570 measurements covering 1,375 solutes and 70 solvents over a temperature range of 243\.15–425\.77 K, following the filtering and deduplication procedures described below\.

Generalization was assessed using two independent external benchmarks, SolProp\[[24](https://arxiv.org/html/2609.21151#bib.bib24)\]covering 6,236 measurements and Leeds\[[6](https://arxiv.org/html/2609.21151#bib.bib6)\]with 1,469 measurements\. Before model training, all solutes present in either benchmark were removed from BigSolDB\. The resulting benchmark evaluations therefore measure performance on compounds that were not encountered during training\.

Transfer learning was evaluated using two aqueous solubility datasets with distinct sizes and chemical compositions\. AqSolDB\[[26](https://arxiv.org/html/2609.21151#bib.bib26)\]which included 9,982 compounds compiled from nine curated sources and spans a broad, heterogeneous region of drug\-like and industrial chemical space\. MoleculeNet ESOL\[[12](https://arxiv.org/html/2609.21151#bib.bib12)\]contains 1,128 compounds and represents a smaller, chemically narrower and more uniformly curated benchmark\. To prevent overlap with the pretraining data, compounds present in the aqueous subset of BigSolDB were removed from each target dataset using InChIKey matching\. This approach captures equivalent molecular representations that may differ in their raw SMILES encoding\. Deduplication removed 232 compounds from AqSolDB, leaving 9,748, and 46 compounds from ESOL, leaving 1,082\. To examine how transfer performance depends on the amount of target data, fixed random subsets of 1,000 compounds were also sampled from the deduplicated AqSolDB dataset\.

### 4\.2Model architecture

EnSol is an environment\-aware molecular solubility prediction framework that jointly encodes solute and solvent molecules and explicitly conditions predictions on temperature\. The model consists of graph\-based molecular encoders, interaction\-aware coupling between solute and solvent representations, continuous temperature conditioning via feature\-wise modulation, and a probabilistic regression head for solubility prediction\.

#### 4\.2\.1Solute and solvent molecular encoders

Both solute and solvent molecules are represented as molecular graphsG=\(V,E\)G=\(V,E\), where nodes correspond to atoms and are featurized using 12\-dimensional atom features encoding atomic number, electronegativity, van der Waals radius, formal charge, aromaticity, hydrogen count, valence, hydrogen donor/acceptor status, and SP/SP2/SP3hybridization and edges correspond to their intramolecular interactions using a 6\-dimensional encoding bond type \(single, double, triple, aromatic\), conjugation, and ring membership\.\[[28](https://arxiv.org/html/2609.21151#bib.bib28),[29](https://arxiv.org/html/2609.21151#bib.bib29)\]Each molecule is encoded using an AttentiveFP\[[26](https://arxiv.org/html/2609.21151#bib.bib26)\]graph neural network, performing iterative message passing to update atom\-level embeddings while learning attention weights over neighboring atoms\.\[[30](https://arxiv.org/html/2609.21151#bib.bib30)\]Given initial node features𝐱v\\mathbf\{x\}\_\{v\}, the AttentiveFP encoder produces a set of node embeddingshv\{h\}\_\{v\}and a graph\-level embeddinghs​o​l​u​t​e\{h\}\_\{solute\}orhs​o​l​v​e​n​t\{h\}\_\{solvent\}via pooling\. The solute is encoded by a deeper AttentiveFP network \(2 layers, 2 timesteps, 256 dimension\) to capture more complex intramolecular interactions, while the solvent is encoded by a shallower network \(1 layer, 1 timestep, 256 dimension\), reflecting the asymmetric role of solvent in determining solubility context\.

#### 4\.2\.2Solute–solvent interaction modeling via cross\-attention

To model interaction\-relevant context between solute and solvent molecules, EnSol employs a 4\-head cross\-attention to couple their learned representations\.\[[31](https://arxiv.org/html/2609.21151#bib.bib31)\]Given solute node embeddingsHs∈ℝNs×256\{H\}\_\{s\}\\in\\mathbb\{R\}^\{\{N\}\_\{s\}\\times 256\}and solvent node embeddingsHv∈ℝNv×256\{H\}\_\{v\}\\in\\mathbb\{R\}^\{\{N\}\_\{v\}\\times 256\}, cross\-attention computes context\-dependent representations by allowing solute features to attend to solvent features and vice versa\. Attention scores are computed as:

Attention⁡\(Q,K,V\)=softmax⁡\(Q​KTdk\)​V\\operatorname\{Attention\}\(Q,K,V\)=\\operatorname\{softmax\}\\\!\\left\(\\frac\{QK^\{T\}\}\{\\sqrt\{d\_\{k\}\}\}\\right\)V
whereQs=Hs​WQ\{Q\}\_\{s\}=\{H\}\_\{s\}\{W\}\_\{Q\}andKv=Hv​WK\{K\}\_\{v\}=\{H\}\_\{v\}\{W\}\_\{K\}\. The resulting attended representations capture interaction\-relevant features conditioned on the specific solute–solvent pair\. These node\-level representations are pooled to produce context\-aware graph embeddings\.

#### 4\.2\.3Temperature encoding and continuous conditioning

Temperature is incorporated as an explicit continuous conditioning variable\. Raw temperature valuesTTare first standardized using training\-set statistics

T^=T−μTσT\\hat\{T\}=\\frac\{T\-\\mu\_\{T\}\}\{\\sigma\_\{T\}\}
The standardized temperature is expanded using a 10\-center radial basis function \(RBF\) expansion,

ϕi\(T\)=exp\[−γrbf\(T^−ci\)2\],i=1,…,10\\phi\_\{i\}\(T\)=\\exp\\\!\\left\[\-\\gamma\_\{\\mathrm\{rbf\}\}\\left\(\\hat\{T\}\-c\_\{i\}\\right\)^\{2\}\\right\],\\qquad i=1,\\ldots,10
where the centersc1,…,c10\{c\}\_\{1\},\\ldots,\{c\}\_\{10\}are placed at the 0\.05–0\.95 quantiles of the standardized training\-temperature distribution, rather than on a fixed symmetric grid, so that RBF resolution matches the density of the training data\. The bandwidth is set adaptively from the resulting center spacing as

γrbf=12​\[median⁡\(diff⁡\(sort⁡\(c\)\)\)\]2\\gamma\_\{\\mathrm\{rbf\}\}=\\frac\{1\}\{2\\left\[\\operatorname\{median\}\\\!\\left\(\\operatorname\{diff\}\(\\operatorname\{sort\}\(c\)\)\\right\)\\right\]^\{2\}\}
As RBF features decay to zero outside the region spanned by the centers, temperatures beyond the training range would otherwise be indistinguishable\. We therefore create two monotonic extrapolation features that grow without saturating, preserving rank order into the tails

overflow⁡\(T^\)=log⁡\[1\+ReLU⁡\(T^−maxi⁡\(ci\)\)\]\\operatorname\{overflow\}\(\\hat\{T\}\)=\\log\\\!\\left\[1\+\\operatorname\{ReLU\}\\\!\\left\(\\hat\{T\}\-\\max\_\{i\}\(c\_\{i\}\)\\right\)\\right\]
underflow⁡\(T^\)=log⁡\[1\+ReLU⁡\(mini⁡\(ci\)−T^\)\]\\operatorname\{underflow\}\(\\hat\{T\}\)=\\log\\\!\\left\[1\+\\operatorname\{ReLU\}\\\!\\left\(\\min\_\{i\}\(c\_\{i\}\)\-\\hat\{T\}\\right\)\\right\]
By appending the overflow and underflow features we end up with the temperature feature vector of

φ⁡\(T\)=\[ϕ1​\(T\),…,ϕ10​\(T\),overflow⁡\(T^\),underflow⁡\(T^\)\]∈ℝ12\\varphi\(T\)=\\left\[\\phi\_\{1\}\(T\),\\ldots,\\phi\_\{10\}\(T\),\\operatorname\{overflow\}\(\\hat\{T\}\),\\operatorname\{underflow\}\(\\hat\{T\}\)\\right\]\\in\\mathbb\{R\}^\{12\}

#### 4\.2\.4Environment embedding via feature\-wise linear modulation

The temperature feature vector is projected through a two\-layer perceptron that emits both scale and shift parameters for feature\-wise linear modulation \(FiLM\):

\[γraw,β\]=W2​ReLU⁡\(W1​φ​\(T\)\)∈ℝ2​d\[\\gamma\_\{\\mathrm\{raw\}\},\\beta\]=W\_\{2\}\\operatorname\{ReLU\}\\\!\\left\(W\_\{1\}\\varphi\(T\)\\right\)\\in\\mathbb\{R\}^\{2d\}
γ=1\+tanh⁡\(γraw\)\\gamma=1\+\\tanh\(\\gamma\_\{\\mathrm\{raw\}\}\)
Given a solvent graph\-level embeddinghs​o​l​v​e​n​t∈ℝd\{h\}\_\{solvent\}\\in\\mathbb\{R\}^\{d\}, the environment embedding is computed as

henv=hsolvent⊙γ\+βh\_\{\\mathrm\{env\}\}=h\_\{\\mathrm\{solvent\}\}\\odot\\gamma\+\\beta
where⊙\\odotdenotes element\-wise multiplication\. Parameterizing the scale asγ=1\+tanh⁡\(γraw\)\\gamma=1\+\\tanh\(\\gamma\_\{\\mathrm\{raw\}\}\)centers the transformation on the identity at initialization and constrainsγ∈\(0,2\)\\gamma\\in\(0,2\), so that temperature can both attenuate and amplify individual dimensions of the solvent embedding, while the additive shiftβ\\betaallows temperature to translate the representation\. The resulting solvent–temperature context vectorhe​n​v\{h\}\_\{env\}captures continuous thermodynamic modulation of solvent behavior\. This full affine formulation replaces a scale\-only variant used in preliminary experiments, in which the temperature embedding was passed through a sigmoid and applied multiplicatively; because that gate was restricted to\(0,1\)\(0,1\)it could only attenuate the solvent embedding and could not represent amplification or translation\.

#### 4\.2\.5Probabilistic solubility prediction via mixture density network

The final solute embedding and environment embedding are concatenated and passed to a MDN implemented as a multilayer perceptron\. The MDN predicts parameters of a mixture ofK=3K=3Gaussian components for solubilityyy

p⁡\(y∣𝐱\)=∑k=1Kπk​\(𝐱\)​𝒩​\(y∣μk​\(𝐱\),σk2​\(𝐱\)\)p\(y\\mid\\mathbf\{x\}\)=\\sum\_\{k=1\}^\{K\}\\pi\_\{k\}\(\\mathbf\{x\}\)\\,\\mathcal\{N\}\\\!\\left\(y\\mid\\mu\_\{k\}\(\\mathbf\{x\}\),\\sigma\_\{k\}^\{2\}\(\\mathbf\{x\}\)\\right\)
whereπk\{\\pi\}\_\{k\}are mixture weights satisfying∑kπk=1\\sum\_\{k\}\{\\pi\}\_\{k\}=1, andμk\{\\mu\}\_\{k\}andσk2\{\\sigma\}\_\{k\}^\{2\}denote component means and variances and N is the number of training samples\. The model is trained by minimizing the negative log\-likelihood

ℒ=−∑n=1Nlog\[∑k=1Kπk\(𝐱n\)𝒩\(yn∣μk\(𝐱n\),σk2\(𝐱n\)\)\]\\mathcal\{L\}=\-\\sum\_\{n=1\}^\{N\}\\log\\\!\\left\[\\sum\_\{k=1\}^\{K\}\\pi\_\{k\}\(\\mathbf\{x\}\_\{n\}\)\\,\\mathcal\{N\}\\\!\\left\(y\_\{n\}\\mid\\mu\_\{k\}\(\\mathbf\{x\}\_\{n\}\),\\sigma\_\{k\}^\{2\}\(\\mathbf\{x\}\_\{n\}\)\\right\)\\right\]
The predictive mean is given by

y^=∑k=1Kπk​μk\\hat\{y\}=\\sum\_\{k=1\}^\{K\}\\pi\_\{k\}\\mu\_\{k\}
and the predictive variance decomposes into aleatoric and mixture\-induced uncertainty

Var⁡\(y\)=∑k=1Kπk​σk2\+∑k=1Kπk​\(μk−y^\)2\\operatorname\{Var\}\(y\)=\\sum\_\{k=1\}^\{K\}\\pi\_\{k\}\\sigma\_\{k\}^\{2\}\+\\sum\_\{k=1\}^\{K\}\\pi\_\{k\}\(\\mu\_\{k\}\-\\hat\{y\}\)^\{2\}
This probabilistic formulation allows EnSol to represent heteroscedastic, non\-Gaussian predictive distributions and to report predictive uncertainty alongside point predictions\.

### 4\.3Model training and optimization

Models were trained using the Adam optimizer with a learning rate of1×10−41\\times 10^\{\-4\}\. Training used a batch size of 32 for a maximum of 10 epochs, with the checkpoint achieving the best validation Spearman correlation retained as the final model\. Training was performed on a single NVIDIA A100\-SXM4\-80GB GPU using PyTorch 2\.7\.1 and PyTorch Geometric 2\.6\.0\. Data were partitioned at the solute level \(grouped by unique solute SMILES, so no solute appeared in more than one fold\) into 90:10 train–validation splits\. All experiments were run with three random seeds, each controlling both model initialization and split composition, so that reported variability reflects both sources\.

### 4\.4Experimental workflows

Solubilities of 10 molecules in 10 solvents were chosen to verify the model\. Chemicals and solvents were purchased from Sigma Aldrich, Ambeed, Chemscene, Thermo Fisher Scientific, Oakwood Chemical and were used without further purification\. Solubility measurements were performed by visual inspection of complete dissolution\. Compounds exhibiting solubilities ¡1 mg mL\-1were considered fully insoluble\. For compounds with intermediate solubilities \(1–10 mg mL\-1\), solvent was titrated into 10 mg of substrate until a clear solution was obtained\. For compounds with solubilities ¿10 mg mL\-1, substrate was added incrementally to 1 mL of solvent until the saturation point was reached, as indicated by the presence of persistent undissolved solid\.

### 4\.5Code Availability

## Acknowledgements

This work was supported by the U\.S\. National Science Foundation \(NSF\) \(CHE\-2505932 \(H\.J\. and H\.Z\.\) and DBI\-2400058 \(H\.Z\.\)\)\. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the author\(s\) and do not necessarily reflect those of the NSF\. Our research benefitted from the computing resources at Delta, the National Center for Supercomputing Applications, enabled by allocation BIO250057 from the Advanced Cyberinfrastructure Coordination Ecosystem: Services & Support \(ACCESS\) program, funded by NSF grants 2138259, 2138286, 2138307, 2137603, and 2138296\.

## Contributions

H\.J\. and H\.Z\. coordinated the project\. T\.N and S\.S jointly designed the research\. T\.N, S\.S and Z\.Z performed the research\. T\.N and S\.S analyzed the data\. T\.N, S\.S and Z\.Z wrote the paper\. All authors approved the final paper\.

## Declarations

### Competing interests

The authors declare no competing interests\.

## References

- \(1\)Barrett, J\. A\., Yang, W\., Skolnik, S\. M\., Belliveau, L\. M\. & Patros, K\. M\. Discovery solubility measurement and assessment of small molecules with drug development in mind\.Drug Discov\. Today27, 1315–1325 \(2022\)\.
- \(2\)Wu, K\.et al\.Overcoming Challenges in Small\-Molecule Drug Bioavailability: A Review of Key Factors and Approaches\.Int\. J\. Mol\. Sci\.25, 13121 \(2024\)\.
- \(3\)Tayyebi, A\.et al\.Prediction of organic compound aqueous solubility using machine learning: a comparison study of descriptor\-based and fingerprints\-based models\.J\. Cheminformatics15, 99 \(2023\)\.
- \(4\)Pendam, D\.et al\.Advances in formulation strategies and stability considerations of amorphous solid dispersions\.J\. Drug Deliv\. Sci\. Technol\.108, 106922 \(2025\)\.
- \(5\)Bloomquist, C\. K\., Zhang, Z\., Aydil, E\. S\. & Modestino, M\. A\. Understanding the effects of multiphase flow properties in transport\-limited organic electrosynthesis\.Chem\. Eng\. J\.528, 172125 \(2026\)\.
- \(6\)Boobier, S\., Hose, D\. R\. J\., Blacker, A\. J\. & Nguyen, B\. N\. Machine learning with physicochemical relationships: solubility prediction in organic solvents and water\.Nat\. Commun\.11, 5753 \(2020\)\.
- \(7\)Murdande, S\. B\., Pikal, M\. J\., Shanker, R\. M\. & Bogner, R\. H\. Aqueous solubility of crystalline and amorphous drugs: Challenges in measurement\.Pharm\. Dev\. Technol\.16, 187–200 \(2011\)\.
- \(8\)Sharma, V\.et al\.Toward microfluidic continuous\-flow and intelligent downstream processing of biopharmaceuticals\.Lab\. Chip24, 2861–2882 \(2024\)\.
- \(9\)Rahimpour, E\., Alvani\-Alamdari, S\., Acree, W\. E\. & Jouyban, A\. Drug Solubility Correlation Using the Jouyban–Acree Model: Effects of Concentration Units and Error Criteria\.Molecules27, 1998 \(2022\)\.
- \(10\)Lipinski, C\. A\., Lombardo, F\., Dominy, B\. W\. & Feeney, P\. J\. Experimental and computational approaches to estimate solubility and permeability in drug discovery and development settings\.Adv\. Drug Deliv\. Rev\.23, 3–25 \(1997\)\.
- \(11\)Könczöl, Á\. & Dargó, G\. Brief overview of solubility methods: Recent trends in equilibrium solubility measurement and predictive models\.Drug Discov\. Today Technol\.27, 3–10 \(2018\)\.
- \(12\)Delaney, John S\. ESOL: estimating aqueous solubility directly from molecular structure\.Journal of chemical information and computer sciences\.44\.3 \(2004\)
- \(13\)Llompart, P\.et al\.Will we ever be able to accurately predict solubility?Sci\. Data11, 303 \(2024\)\.
- \(14\)Jouyban, A\., Rahimpour, E\. & Karimzadeh, Z\. A new correlative model to simulate the solubility of drugs in mono\-solvent systems at various temperatures\.J\. Mol\. Liq\.343, 117587 \(2021\)\.
- \(15\)Al Ibrahim, E\., Morgan, N\., Müller, S\., Motati, S\. & Green, W\. H\. Accurately Predicting Solubility Curves via a Thermodynamic Cycle, Machine Learning, and Solvent Ensembles\.J\. Am\. Chem\. Soc\.147, 45057–45069 \(2025\)\.
- \(16\)Panapitiya, G\.et al\.Evaluation of Deep Learning Architectures for Aqueous Solubility Prediction\.ACS Omega7, 15695–15710 \(2022\)\.
- \(17\)Francoeur, P\. G\. & Koes, D\. R\. SolTranNet–A Machine Learning Tool for Fast Aqueous Solubility Prediction\.J\. Chem\. Inf\. Model\.61, 2530–2536 \(2021\)\.
- \(18\)Ghanavati, M\. A\., Ahmadi, S\. & Rohani, S\. A machine learning approach for the prediction of aqueous solubility of pharmaceuticals: a comparative model and dataset analysis\.Digit\. Discov\.3, 2085–2104 \(2024\)\.
- \(19\)Attia, L\., Burns, J\. W\., Doyle, P\. S\. & Green, W\. H\. Data\-driven organic solubility prediction at the limit of aleatoric uncertainty\.Nat\. Commun\.16, 7497 \(2025\)\.
- \(20\)Gao, P\.et al\.Accurate predictions of drugs aqueous solubility via deep learning tools\.J\. Mol\. Struct\.1249, 131562 \(2022\)\.
- \(21\)Perez, E\., Strub, F\., de Vries, H\., Dumoulin, V\. & Courville, A\. FiLM: Visual Reasoning with a General Conditioning Layer\. Preprint at[https://doi\.org/10\.48550/ARXIV\.1709\.07871](https://doi.org/10.48550/ARXIV.1709.07871)\(2017\)\.
- \(22\)Bishop, C\. M\.Mixture Density Networks\. \(1994\)\.
- \(23\)Krasnov, L\.et al\.BigSolDB 2\.0, dataset of solubility values for organic compounds in different solvents at various temperatures\.Sci\. Data12, 1236 \(2025\)\.
- \(24\)Vermeire, F\. H\., Chung, Y\. & Green, W\. H\. Predicting Solubility Limits of Organic Solutes for a Wide Range of Solvents and Temperatures\.J\. Am\. Chem\. Soc\.144, 10785–10797 \(2022\)\.
- \(25\)Ottoboni, S\.et al\.A Novel Integrated Workflow for Isolation Solvent Selection Using Prediction and Modeling\.Org\. Process Res\. Dev\.25, 1143–1159 \(2021\)\.
- \(26\)Sorkun, M\.C\., Khetan, A\. & Er, S\. AqSolDB, a curated reference set of aqueous solubility and 2D descriptors for a diverse set of compounds\.Sci Data6, 143 \(2019\)\.
- \(27\)Wang, Z\.et al\.Advanced graph and sequence neural networks for molecular property prediction and drug discovery\.Bioinformatics38, 2579–2586 \(2022\)\.
- \(28\)Zhao, B\., Xu, W\., Guan, J\. & Zhou, S\. Molecular property prediction based on graph structure learning\.Bioinformatics40, btae304 \(2024\)\.
- \(29\)Fang, X\.et al\.Geometry\-enhanced molecular representation learning for property prediction\.Nat\. Mach\. Intell\.4, 127–134 \(2022\)\.
- \(30\)Xiong, Z\.et al\.Pushing the Boundaries of Molecular Representation for Drug Discovery with the Graph Attention Mechanism\.J\. Med\. Chem\.63, 8749–8760 \(2020\)\.
- \(31\)Vaswani, A\.et al\.Attention Is All You Need\. Preprint at[https://doi\.org/10\.48550/ARXIV\.1706\.03762](https://doi.org/10.48550/ARXIV.1706.03762)\(2017\)\.

相似文章

利用基于图的工具改进小型语言模型中的分子性质预测

arXiv cs.AI

本文提出了一种上下文增强提示框架,利用图神经网络专家模型提供预测提示和解释性子图,以改进小型语言模型中的分子性质预测。在MUTAG和Tox21上的实验显示,与仅使用SMILES的基线相比,准确率提升高达74%。