SMILESGNN: 通过SMILES-图交叉注意力融合的可解释临床毒性预测

arXiv cs.LG 论文

摘要

SMILESGNN引入了一种多模态架构,该架构结合了SMILES转换器和图神经网络,并采用交叉注意力机制,用于可解释的药物毒性预测。它在ClinTox和Tox21等基准测试上以最少的参数取得了有竞争力的性能。

arXiv:2609.28553v1 Announce Type: new Abstract: Drug toxicity prediction is critical for reducing late-stage attrition in drug discovery, yet remains challenging due to severe class imbalance, scaffold-based generalization, and the clinical need for interpretable predictions. Single-modality approaches-SMILES Transformers or graph neural networks capture complementary aspects of molecular structure, while sequence-only models cannot directly provide graph-attributed explanations. We present SMILESGNN, a multimodal architecture that fuses a SMILES Transformer encoder and a GATv2 graph encoder via cross-attention, and SMILESGNN-PT, a variant using a ChemBERTa-2 pretrained backbone. The design retains an explicit graph branch within the predictive pipeline, supporting GNNExplainer-based analysis of substructures associated with toxic predictions. On ClinTox, SMILESGNN achieves AUC-ROC 0.987 and F1 0.906 with only 0.4M parameters, performing competitively with a strong SMILESTransformer and a larger ChemBERTa-2/GATv2 concat-fusion baseline. On Tox21 (12 tasks), SMILESGNN-PT obtains mean AUC-ROC 0.750, comparable to ChemBERTa-2 alone and the same-backbone concat-fusion baseline. Overall, the results suggest that cross-attention is a practical fusion alternative that preserves competitive predictive performance while enabling graph-based interpretability support.
查看原文
查看缓存全文

缓存时间: 2026/09/25 09:26

# SMILESGNN: Interpretable Clinical Toxicity Prediction via SMILES-Graph Cross-Attention Fusion
Source: [https://arxiv.org/html/2609.28553](https://arxiv.org/html/2609.28553)
Thuy Quynh NguyenAffiliation:National Economics University Hanoi, Vietnam 11247346@st\.neu\.edu\.vnDuc Minh LeAffiliation:National Economics University Hanoi, Vietnam 11247320@st\.neu\.edu\.vnAffiliation:\[ \]Ho Nhat Minh NguyenAffiliation:National Economics University Hanoi, Vietnam 11247321@st\.neu\.edu\.vnThanh Long Dai DoanAffiliation:VNU University of Science Hanoi, Vietnam doandaithanhlong\_sdh21@hus\.edu\.vnTrong Nghia Nguyen††thanks:Corresponding author: nghiant@neu\.edu\.vnAffiliation:National Economics University Hanoi, Vietnam nghiant@neu\.edu\.vn

###### Abstract

Drug toxicity prediction is critical for reducing late\-stage attrition in drug discovery, yet remains challenging due to severe class imbalance, scaffold\-based generalization, and the clinical need for interpretable predictions\. Single\-modality approaches\-SMILES Transformers or graph neural networks\-capture complementary aspects of molecular structure, while sequence\-only models cannot directly provide graph\-attributed explanations\. We present SMILESGNN, a multimodal architecture that fuses a SMILES Transformer encoder and a GATv2 graph encoder via cross\-attention, and SMILESGNN\-PT, a variant using a ChemBERTa\-2 pretrained backbone\. The design retains an explicit graph branch within the predictive pipeline, supporting GNNExplainer\-based analysis of substructures associated with toxic predictions\. On ClinTox, SMILESGNN achieves AUC\-ROC 0\.987 ± 0\.012 and F1 0\.906 ± 0\.039 with only 0\.4M parameters, performing competitively with a strong SMILES Transformer and a larger ChemBERTa\-2/GATv2 concat\-fusion baseline\. On Tox21 \(12 tasks\), SMILESGNN\-PT obtains mean AUC\-ROC 0\.750 ± 0\.002, comparable to ChemBERTa\-2 alone and the same\-backbone concat\-fusion baseline\. Overall, the results suggest that cross\-attention is a practical fusion alternative that preserves competitive predictive performance while enabling graph\-based interpretability support\.

###### Index Terms:

drug toxicity prediction, graph neural network, multimodal fusion, cross\-attention

## IIntroduction

Drug discovery is a lengthy and costly process, with clinical development alone spanning over a decade\[[1](https://arxiv.org/html/2609.28553#bib.bib16)\]\. A primary driver of late\-stage attrition is unexpected toxicity: compounds that pass early screening may still fail in clinical trials due to organ toxicity or adverse drug reactions not apparent until human exposure\. Early toxicity prediction from molecular structure is therefore critical, enabling researchers to prioritize safe candidates and reduce costly failures\.

Three core challenges make clinical toxicity prediction particularly difficult\. First,class imbalance: the ClinTox benchmark\[[2](https://arxiv.org/html/2609.28553#bib.bib1)\]has a 1:11\.5 toxic\-to\-nontoxic ratio, causing models to collapse toward the majority class\. Second,scaffold\-based generalization: models must predict toxicity for structurally novel compounds unseen during training, making scaffold\-based evaluation essential\. Third,multi\-scale structural complexity: toxicity arises from local features \(reactive functional groups, halogenated rings\) and global scaffold geometry simultaneously\.

Deep learning has enabled powerful molecular representations, yet single\-modality approaches — SMILES Transformers or GNNs — each capture only part of the structural picture\. This motivates multimodal fusion\. Beyond predictive accuracy,interpretabilityis a distinct clinical requirement: toxicologists need to identify which atomic substructures drive a toxic classification\. Sequence\-only models cannot provide graph\-attributed explanations; GNNExplainer\[[3](https://arxiv.org/html/2609.28553#bib.bib21)\]requires an explicit graph pathway in the predictive model\. Cross\-attention fusion is intended to couple SMILES and graph representations so that post\-hoc graph attributions can be analyzed for the GATv2 branch, while ablation results quantify how much each branch contributes\. A further open question is whether cross\-attention outperformsconcatenation— none of the existing methods address either requirement under the extreme class imbalance of clinical toxicity prediction\.

To address these gaps, we propose SMILESGNN, a multimodal architecture fusing a SMILES Transformer encoder and a GATv2 graph encoder via cross\-attention, trained with focal loss for class imbalance, and SMILESGNN\-PT, its pretrained\-backbone variant using ChemBERTa\-2\[[4](https://arxiv.org/html/2609.28553#bib.bib19)\]\. Our contributions are: \(1\) the dual\-pathway cross\-attention architecture enablesGNNExplainer\-based graph attribution, highlighting chemically interpretable toxicophore\-like substructures\-a capability sequence\-only models cannot provide; \(2\) SMILESGNN achieves competitive AUC\-ROC \(±0\.0120\.987\\\!\\pm\\\!0\.012\), outperforms all graph\-based single\-modality baselines, and exceeds the sequence\-only baseline in minority\-class F1 with only 0\.4M parameters; \(3\) a controlled comparison under identical backbone conditions shows cross\-attention is competitive with concatenation while retaining an explicit graph branch for interpretability; \(4\) SMILESGNN\-PT generalizes to pretrained backbones, matching the same\-backbone concat\-fusion baseline on both ClinTox and Tox21 \(12 tasks\)\.

## IIRelated Work

Molecular fingerprints such as ECFP\[[5](https://arxiv.org/html/2609.28553#bib.bib7)\]remain widely used baselines but lose structural detail through fixed\-length hashing\. Graph Neural Networks \(GNNs\) operate directly on molecular graphs via message passing; prominent variants include GIN\[[6](https://arxiv.org/html/2609.28553#bib.bib9)\], GATv2\[[7](https://arxiv.org/html/2609.28553#bib.bib3)\], DMPNN\[[8](https://arxiv.org/html/2609.28553#bib.bib10)\], AttentiveFP\[[9](https://arxiv.org/html/2609.28553#bib.bib11)\], and the self\-supervised pretrained GNN backbone of\[[10](https://arxiv.org/html/2609.28553#bib.bib8)\]\. On the sequence side, Transformer models\[[11](https://arxiv.org/html/2609.28553#bib.bib12)\]treat SMILES strings as token sequences, with ChemBERTa\-2\[[4](https://arxiv.org/html/2609.28553#bib.bib19)\]and MoLFormer\-XL\[[12](https://arxiv.org/html/2609.28553#bib.bib18)\]demonstrating strong performance through large\-scale SMILES pretraining\[[13](https://arxiv.org/html/2609.28553#bib.bib17)\]\.

Multimodal methods combine these complementary views\. FP\-GNN\[[14](https://arxiv.org/html/2609.28553#bib.bib22)\]showed that concatenating ECFP fingerprints with GNN embeddings consistently outperforms either modality alone\. More recently, Zang et al\.\[[15](https://arxiv.org/html/2609.28553#bib.bib20)\]proposed gated fusion with representation decoupling, showing selective modality weighting outperforms fixed concatenation\. However, none of these methods address extreme class imbalance, provide a controlled cross\-attention versus concatenation comparison, or evaluate graph\-aware post\-hoc attribution under a multimodal toxicity setting\.

## IIIMethodology

![Refer to caption](https://arxiv.org/html/2609.28553v1/Molecule_method.jpg)

Fig\. 1:SMILESGNN architecture: \(1\) Transformer SMILES encoder; \(2\) GATv2 graph encoder; \(3\) cross\-attention fusion module; \(4\) MLP predictor head\. The dual\-pathway design captures complementary sequence and graph structural information; cross\-attention couples the GATv2 pathway with the SMILES pathway, supporting GNNExplainer\-based graph interpretability\.### III\-ADataset

ClinTox\[[2](https://arxiv.org/html/2609.28553#bib.bib1)\]comprises 1,480 drug\-like molecules with binary clinical toxicity labels \(CT\_TOX task\), exhibiting pronounced class imbalance \(11\.5:1\)\. We apply scaffold\-based splitting\[[16](https://arxiv.org/html/2609.28553#bib.bib2)\]\(seed 42, 80/10/10\) to evaluate generalization to novel molecular scaffolds\. Fig\.[2](https://arxiv.org/html/2609.28553#S3.F2)illustrates that structurally similar molecules can have markedly different clinical outcomes\.

![Refer to caption](https://arxiv.org/html/2609.28553v1/non_toxic_example_1.png)\(a\)Non\-toxic \(FDA\-approved\)
![Refer to caption](https://arxiv.org/html/2609.28553v1/toxic_example_0.png)\(b\)Toxic \(failed clinical trials\)

Fig\. 2:ClinTox examples: structurally similar molecules with opposite clinical outcomes\.We additionally evaluate on Tox21\[[2](https://arxiv.org/html/2609.28553#bib.bib1)\], comprising 7,831 molecules across 12 binary assays \(nuclear receptors and stress response pathways\) with up to 30% missing annotations per task\-handled via masked focal loss\. The same scaffold split \(seed 42, 80/10/10\) yields 6,264/783/784 training/validation/test molecules; performance is mean AUC\-ROC and mean PR\-AUC across all 12 tasks\.

### III\-BData Processing

#### III\-B1SMILES Sequence Processing

For SMILESGNN, a custom tokenizer decomposes SMILES into chemical units \(vocabulary 69, derived from the ClinTox training set\[[17](https://arxiv.org/html/2609.28553#bib.bib13)\]\)\. For SMILESGNN\-PT, ChemBERTa\-2 WordPiece subword tokenization is used\. Both variants pad or truncate to 128 tokens\.

#### III\-B2Molecular Graph Processing

Atom nodes use 25 features \(atomic number, formal charge, hybridization, stereochemistry, ring membership, aromaticity\); edges use 17 features \(bond type, direction, ring membership, conjugation, stereochemistry\)\.

### III\-CSMILESGNN Architecture

The overall architecture is illustrated in Fig\.[1](https://arxiv.org/html/2609.28553#S3.F1)\.

#### III\-C1SMILES Encoder

The SMILES encoder employs a Transformer architecture with embedding dimensiondmodel=96d\_\{\\text\{model\}\}=96and vocabulary size 69\. Token embeddings are followed by learnable positional encodings\. Two Transformer encoder layers, each with 4 attention heads and feedforward dimension 192, process the sequences with dropout 0\.4\. Masked mean pooling aggregates the sequence into𝐡SMILES∈ℝ96\\mathbf\{h\}\_\{\\text\{SMILES\}\}\\in\\mathbb\{R\}^\{96\}\.

#### III\-C2Graph Encoder

The graph encoder uses GATv2\[[7](https://arxiv.org/html/2609.28553#bib.bib3)\]\. Atom features \(25 dimensions\) are projected to hidden dimension 96; edge features are embedded when available\. Three GATv2 layers with 4 attention heads perform message passing with LayerNorm, ReLU, residual connections, and dropout 0\.4\. Jumping Knowledge connections\[[18](https://arxiv.org/html/2609.28553#bib.bib4)\]concatenate all layer outputs to produce 288\-dim node representations\. Mean\-Max pooling generates𝐡graph∈ℝ576\\mathbf\{h\}\_\{\\text\{graph\}\}\\in\\mathbb\{R\}^\{576\}\.

#### III\-C3Attention\-Based Fusion Module

The fusion module combines SMILES and graph representations through multi\-head cross\-attention\[[19](https://arxiv.org/html/2609.28553#bib.bib15)\]with 4 heads\. The graph representation \(576 dimensions\) is projected to match the SMILES dimension \(96\); the SMILES representation serves as query, the projected graph as key and value\. The attended output is concatenated with the original SMILES representation, yielding𝐡fused∈ℝ192\\mathbf\{h\}\_\{\\text\{fused\}\}\\in\\mathbb\{R\}^\{192\}, allowing the fused representation to condition sequence features on graph features\.

#### III\-C4Predictor Head

A four\-layer MLP maps the fused 192\-dimensional representation to a binary toxicity prediction:192→192→96→48→1192\\rightarrow 192\\rightarrow 96\\rightarrow 48\\rightarrow 1\. Each hidden layer includes BatchNorm, ReLU, and dropout \(0\.4\)\.

### III\-DSMILESGNN\-PT Architecture

SMILESGNN\-PT replaces the custom Transformer SMILES encoder with a ChemBERTa\-2\[[4](https://arxiv.org/html/2609.28553#bib.bib19)\]pretrained backbone \(RoBERTa pretrained on 77M PubChem SMILES, hidden size 384\), while the GATv2 graph encoder is kept identical\. The graph representation \(576 dimensions\) is projected to 384 via a linear layer with LayerNorm\. Cross\-attention fusion operates at 384 dimensions \(ChemBERTa\-2 mean\-pool as query, projected graph as key/value\), yielding𝐡fused∈ℝ768\\mathbf\{h\}\_\{\\text\{fused\}\}\\in\\mathbb\{R\}^\{768\}by concatenation\. The MLP predictor maps→→→→output\_dim768\\\!\\rightarrow\\\!384\\\!\\rightarrow\\\!192\\\!\\rightarrow\\\!96\\\!\\rightarrow\\\!\\text\{output\\\_dim\}\. SMILESGNN\-PT has 4\.7M total parameters \(backbone 3\.4M, head 1\.3M\), trained with a two\-group AdamW \(lrbb=10−4\\text\{lr\}\_\{\\text\{bb\}\}\\\!=\\\!10^\{\-4\},lrhd=10−3\\text\{lr\}\_\{\\text\{hd\}\}\\\!=\\\!10^\{\-3\}\), linear warmup, and gradient clipping \(max norm 1\.0\)\.

### III\-ETraining Methodology

#### III\-E1Loss Function

To address severe class imbalance, we employ Focal Loss\[[20](https://arxiv.org/html/2609.28553#bib.bib5)\]:

ℒfocal=−α​\(1−pt\)γ​log⁡\(pt\)\\mathcal\{L\}\_\{\\text\{focal\}\}=\-\\alpha\(1\-p\_\{t\}\)^\{\\gamma\}\\log\(p\_\{t\}\)\(1\)whereptp\_\{t\}is the predicted probability for the true class,α=0\.25\\alpha=0\.25, andγ=2\.0\\gamma=2\.0\. For ClinTox \(binary\), this loss is applied directly\. For Tox21 \(multi\-task\), we use a masked variant that computes focal loss only over labelled \(compound, assay\) pairs, ignoring missing annotations\.

#### III\-E2Optimization and Regularization

SMILESGNN uses AdamW\[[21](https://arxiv.org/html/2609.28553#bib.bib6)\]\(lr=×10−4\\text\{lr\}=5\\\!\\times\\\!10^\{\-4\}, weight decay10−410^\{\-4\}\), batch size 32, up to 100 \(ClinTox\) or 60 \(Tox21\) epochs, early stopping on validation F1 / mean AUC\-ROC \(patience 20 / 10\)\. SMILESGNN\-PT uses the two\-group AdamW from Section[III\-D](https://arxiv.org/html/2609.28553#S3.SS4), batch size 16, otherwise identical\. A weighted sampler, dropout \(0\.3\-0\.4\), and batch normalization provide regularization\.

#### III\-E3Evaluation Metrics

Models are evaluated on the held\-out test set using AUC\-ROC, AUPRC, F1, and accuracy\.

## IVExperiments and Results

We evaluate SMILESGNN and SMILESGNN\-PT against eight baselines on ClinTox \(scaffold split, 1,184 train / 148 test\)\. Single\-modality baselines: Baseline MLP\[[5](https://arxiv.org/html/2609.28553#bib.bib7)\], Pretrained GIN\[[10](https://arxiv.org/html/2609.28553#bib.bib8)\], GIN\[[6](https://arxiv.org/html/2609.28553#bib.bib9)\], GATv2∗\[[7](https://arxiv.org/html/2609.28553#bib.bib3)\], DMPNN\[[8](https://arxiv.org/html/2609.28553#bib.bib10)\], AttentiveFP\[[9](https://arxiv.org/html/2609.28553#bib.bib11)\], and SMILESTransformer∗\[[11](https://arxiv.org/html/2609.28553#bib.bib12)\]\. Multimodal concat baseline: CB2\-GATv2†\(see Table[I](https://arxiv.org/html/2609.28553#S4.T1)caption\), trained under identical conditions\. SMILESGNN\-PT‡\(Section[III\-D](https://arxiv.org/html/2609.28553#S3.SS4)\) uses the same ChemBERTa\-2 backbone as CB2\-GATv2†with cross\-attention instead of concatenation, directly isolating the fusion strategy\. Models marked∗serve as single\-modality ablation components\.

### IV\-AOverall Performance

TABLE I:Performance on ClinTox test set \(scaffold split, 148 samples\)\.Bold: proposed method\.∗Single\-modality ablation components of SMILESGNN\.†Concat\-fusion baseline: CB2\-GATv2 pairs ChemBERTa\-2\[[4](https://arxiv.org/html/2609.28553#bib.bib19)\]\(3\.4M params\) with GATv2, trained under identical conditions\.‡SMILESGNN\-PT: same ChemBERTa\-2 backbone as CB2\-GATv2 with cross\-attention fusion \(4\.7M params total\)\. All models trained under identical conditions \(focal loss, balanced sampling\)\. Values: mean±\\pmstd, 3 seeds\.Table[I](https://arxiv.org/html/2609.28553#S4.T1)shows that when all models are trained under identical conditions, the SMILES encoder alone \(SMILESTransformer∗,±0\.0030\.995\\\!\\pm\\\!0\.003AUC\-ROC\) is a strong single\-modality baseline\. All five graph\-based baselines are substantially weaker \(best: DMPNN±0\.0110\.893\\\!\\pm\\\!0\.011\), confirming the challenge of scaffold\-based graph generalization on ClinTox\.

SMILESGNN achieves F1±0\.0390\.906\\\!\\pm\\\!0\.039, exceeding SMILESTransformer∗\(F1±0\.0530\.896\\\!\\pm\\\!0\.053\) in minority\-class identification while reducing variance, and matching CB2\-GATv2†\(F1±0\.0630\.890\\\!\\pm\\\!0\.063\)\. In AUC\-ROC \(±0\.0120\.987\\\!\\pm\\\!0\.012\), SMILESGNN is competitive with SMILESTransformer∗within two standard deviations; the graph pathway adds GNNExplainer interpretability rather than a large accuracy gain on this small dataset \(1,184 training molecules\)\.

Replacing concat fusion with cross\-attention under the identical ChemBERTa\-2 backbone yields AUC\-ROC±0\.0140\.979\\\!\\pm\\\!0\.014\-within 0\.008 of CB2\-GATv2†\(±0\.0180\.987\\\!\\pm\\\!0\.018\), with overlapping confidence intervals\-confirming cross\-attention is competitive\. SMILESGNN outperforms SMILESGNN\-PT on ClinTox \(limited data favours the lightweight custom encoder\); this advantage reverses on Tox21’s 6,264 training samples \(\+0\.030\+0\.030AUC\-ROC for SMILESGNN\-PT\)\.

![Refer to caption](https://arxiv.org/html/2609.28553v1/roc_curves_all_models.png)\(a\)ROC Curves
![Refer to caption](https://arxiv.org/html/2609.28553v1/pr_curves_all_models.png)\(b\)Precision\-Recall Curves

Fig\. 3:Performance curves on the ClinTox test set \(seed 42\) for representative models: single\-modality baselines \(AttentiveFP, Pretrained GIN\), multimodal concat baseline \(CB2\-GATv2†\), and the proposed SMILESGNN\. PR curves are more informative for imbalanced datasets; SMILESGNN \(AUPRC 0\.967 on this run\) outperforms all graph\-based baselines and matches the concat\-fusion baseline\.Fig\.[3](https://arxiv.org/html/2609.28553#S4.F3)shows ROC and Precision\-Recall curves for representative models \(seed 42\)\. PR curves confirm SMILESGNN \(mean AUPRC±0\.0280\.946\\\!\\pm\\\!0\.028\) substantially outperforms graph\-based single\-modality baselines, validating multimodal fusion for imbalanced toxicity prediction\[[22](https://arxiv.org/html/2609.28553#bib.bib14)\]\.

We compare SMILESGNN with SMILESTransformer to isolate the graph contribution\. Both models agree on 142 out of 148 test samples; for 3 disagreements favouring SMILESGNN \(Fig\.[4](https://arxiv.org/html/2609.28553#S4.F4)\), it correctly identifies 1 missed toxic compound and avoids 2 false\-positive errors, demonstrating that the graph encoder captures structural signals not apparent from the SMILES sequence alone\.

![Refer to caption](https://arxiv.org/html/2609.28553v1/top2_comparison_smilesgnn_wins.png)

Fig\. 4:Three compounds correctly classified by SMILESGNN but misclassified by SMILESTransformer: one missed toxic detection \(red border\) and two corrected false positives \(green border\), illustrating graph encoder contributions beyond the SMILES sequence\.
### IV\-BFusion Mechanism Ablation

To quantify the contribution of each encoder pathway, we compare three configurations using the same backbone components and training conditions\.

TABLE II:Encoder contribution ablation on ClinTox\. Equal\-capacity projection \(→96576\\\!\\to\\\!96\) applied before cross\-attention fusion\.∗Same encoders as SMILESGNN\.Bold: proposed method\. Values: mean±\\pmstd, 3 seeds\.Table[II](https://arxiv.org/html/2609.28553#S4.T2)reveals two findings: \(1\) GATv2 alone is limited and high\-variance \(F1±0\.1620\.424\\\!\\pm\\\!0\.162\), while SMILESTransformer alone is strong \(F1±0\.0530\.896\\\!\\pm\\\!0\.053, AUC\-ROC±0\.0030\.995\\\!\\pm\\\!0\.003\); \(2\) cross\-attention fusion achieves F1±0\.0390\.906\\\!\\pm\\\!0\.039\-slightly higher than SMILESTransformer alone with lower variance\-suggesting that the graph pathway contributes complementary structural signals for minority\-class identification\. Cross\-attention fusion additionally retains an explicit graph branch for GNNExplainer analysis \(Section[IV\-D](https://arxiv.org/html/2609.28553#S4.SS4)\)\.

### IV\-CTox21 Multi\-Task Evaluation

TABLE III:Mean AUC\-ROC and PR\-AUC on Tox21 \(12 tasks, scaffold split\)\.†Same conditions as Table[I](https://arxiv.org/html/2609.28553#S4.T1)\.‡ChemBERTa\-2 \+ cross\-attn \+ GATv2 \(4\.7M\)\.Bold: best proposed\. Values are mean±\\pmstd over 3 seeds\.On Tox21, SMILESGNN\-PT‡\(±0\.0020\.750\\\!\\pm\\\!0\.002\) outperforms SMILESGNN \(±0\.0080\.720\\\!\\pm\\\!0\.008\) by\+0\.030\+0\.030, confirming the benefit of pretrained representations\. SMILESGNN\-PT‡matches ChemBERTa\-2 alone \(±0\.0090\.752\\\!\\pm\\\!0\.009\) and CB2\-GATv2†\(±0\.0120\.751\\\!\\pm\\\!0\.012\)\-overlapping confidence intervals\-directly demonstrating cross\-attention is competitive with concatenation under the same backbone\.

### IV\-DGraph Representation Analysis

![Refer to caption](https://arxiv.org/html/2609.28553v1/gnnexplainer_interpretability.png)

Fig\. 5:GNNExplainer\[[3](https://arxiv.org/html/2609.28553#bib.bib21)\]interpretability analysis of the six highest\-confidence ClinTox toxic predictions \(Ptox≥0\.954P\_\{\\text\{tox\}\}\\\!\\geq\\\!0\.954\)\. Two blocks of three molecules each;Row 1 \(Atom importance\)andRow 2 \(Bond importance\)per block; darker red indicates higher learned mask importance for the toxic classification\. SMILES embeddings are frozen to isolate graph\-attributed explanations\.To demonstrate model interpretability, we apply GNNExplainer\[[3](https://arxiv.org/html/2609.28553#bib.bib21)\]to the six highest\-confidence toxic predictions from the ClinTox test set \(Fig\.[5](https://arxiv.org/html/2609.28553#S4.F5)\)\. GNNExplainer optimises per\-atom and per\-bond soft masks over the GATv2 pathway to identify compact subgraphs associated with the toxic classification\. Because the graph encoder remains part of the fused predictor, these masks provide a graph\-branch view of the learned decision process; we interpret them as qualitative, model\-specific evidence rather than a formal proof of causality\.

Three chemically interpretable patterns recur across all six molecules\. For halogenated aromatics \(a, d\), atom and bond importance concentrates on halogenated phenyl rings and conjugated bridges, consistent with cytochrome P450\-mediated activation — a well\-documented cytotoxicity mechanism\. For reactive heterocycles \(c, e\), importance peaks at triazole and isoxazole rings, corresponding to structural alerts for metabolic nitrile hydrolysis and ring\-opening electrophilic reactivity\. For extended scaffolds \(b, f\), importance is more diffuse but consistently localises at central amide linkers and terminal halogenated rings\.

## VConclusion

This paper introduced SMILESGNN, a multimodal toxicity prediction framework that combines SMILES sequence modeling, graph message passing, and cross\-attention fusion under class\-imbalance\-aware training\. Its central contribution is to retain an explicit molecular graph pathway while preserving competitive predictive performance, allowing benchmark evaluation to be complemented by GNNExplainer\-based inspection of toxicophore\-like substructures\. The pretrained variant, SMILESGNN\-PT, further suggests that the same fusion strategy can be transferred to a ChemBERTa\-2 backbone and remain competitive with a comparable concat\-fusion design\. However, ClinTox is small and highly imbalanced, the explanations are qualitative rather than causal, and the current evaluation lacks stronger stress tests such as graph perturbation, calibration analysis, and external clinical validation\. Future work will therefore focus on validating the graph branch more rigorously through perturbation\-based ablations, graph\-level pretraining, uncertainty calibration, and evaluation on larger toxicity benchmarks and external clinical datasets\. We also plan to study task\-adaptive loss weighting and explanation stability so that cross\-attention fusion can better support both robust prediction and chemically meaningful interpretation\.

## References

- \[1\]S\. Seal, M\. Mahale, M\. García\-Ortegón, C\. K\. Joshi, L\. Hosseini\-Gerami, A\. Beatson, M\. Greenig, M\. Shekhar, A\. Patra, C\. Weis,et al\.\(2025\)Machine learning for toxicity prediction using chemical structures: pillars for success in the real world\.Chemical research in toxicology38\(5\),pp\. 759–807\.Cited by:[§I](https://arxiv.org/html/2609.28553#S1.p1.1)\.
- \[2\]Z\. Wu, B\. Ramsundar, E\. N\. Feinberg, J\. Gomes, C\. Geniesse, A\. S\. Pappu, K\. Leswing, and V\. Pande\(2018\)MoleculeNet: a benchmark for molecular machine learning\.Chemical Science9\(2\),pp\. 513–530\.External Links:[Document](https://dx.doi.org/10.1039/C7SC02664A)Cited by:[§I](https://arxiv.org/html/2609.28553#S1.p2.1),[§III\-A](https://arxiv.org/html/2609.28553#S3.SS1.p1.1),[§III\-A](https://arxiv.org/html/2609.28553#S3.SS1.p2.1)\.
- \[3\]Z\. Ying, D\. Bourgeois, J\. You, M\. Zitnik, and J\. Leskovec\(2019\)GNNExplainer: generating explanations for graph neural networks\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§I](https://arxiv.org/html/2609.28553#S1.p3.1),[Fig\. 5](https://arxiv.org/html/2609.28553#S4.F5),[§IV\-D](https://arxiv.org/html/2609.28553#S4.SS4.p1.1)\.
- \[4\]W\. Ahmad, E\. Simon, S\. Chithrananda, G\. Grand, and B\. Ramsundar\(2022\)ChemBERTa\-2: towards chemical foundation models\.InNeurIPS 2022 Workshop on AI for Science,Cited by:[§I](https://arxiv.org/html/2609.28553#S1.p4.1),[§II](https://arxiv.org/html/2609.28553#S2.p1.1),[§III\-D](https://arxiv.org/html/2609.28553#S3.SS4.p1.1),[TABLE I](https://arxiv.org/html/2609.28553#S4.T1),[TABLE I](https://arxiv.org/html/2609.28553#S4.T1.10.11.1.1),[TABLE III](https://arxiv.org/html/2609.28553#S4.T3.8.1.5.1),[TABLE III](https://arxiv.org/html/2609.28553#S4.T3.8.1.7.1)\.
- \[5\]D\. Rogers and M\. Hahn\(2010\)Extended\-connectivity fingerprints\.Journal of Chemical Information and Modeling50\(5\),pp\. 742–754\.External Links:[Document](https://dx.doi.org/10.1021/ci100050t)Cited by:[§II](https://arxiv.org/html/2609.28553#S2.p1.1),[TABLE I](https://arxiv.org/html/2609.28553#S4.T1.10.4.1.1),[TABLE III](https://arxiv.org/html/2609.28553#S4.T3.8.1.3.1),[§IV](https://arxiv.org/html/2609.28553#S4.p1.1)\.
- \[6\]K\. Xu, W\. Hu, J\. Leskovec, and S\. Jegelka\(2019\)How powerful are graph neural networks?\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§II](https://arxiv.org/html/2609.28553#S2.p1.1),[TABLE I](https://arxiv.org/html/2609.28553#S4.T1.10.5.1.1),[§IV](https://arxiv.org/html/2609.28553#S4.p1.1)\.
- \[7\]S\. Brody, U\. Alon, and E\. Yahav\(2022\)How attentive are graph attention networks?\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§II](https://arxiv.org/html/2609.28553#S2.p1.1),[§III\-C2](https://arxiv.org/html/2609.28553#S3.SS3.SSS2.p1.1),[TABLE I](https://arxiv.org/html/2609.28553#S4.T1.10.6.1.1),[§IV](https://arxiv.org/html/2609.28553#S4.p1.1)\.
- \[8\]K\. Yang, K\. Swanson, W\. Jin, C\. Coley, P\. Eiden, H\. Gao, A\. Guzman\-Perez, T\. Hopper, B\. Kelley, M\. Mathea, A\. Palmer, V\. Settels, T\. Jaakkola, K\. F\. Jensen, and R\. Barzilay\(2019\)Analyzing learned molecular representations for property prediction\.Journal of Chemical Information and Modeling59\(8\),pp\. 3370–3388\.External Links:[Document](https://dx.doi.org/10.1021/acs.jcim.9b00237)Cited by:[§II](https://arxiv.org/html/2609.28553#S2.p1.1),[TABLE I](https://arxiv.org/html/2609.28553#S4.T1.10.8.1.1),[§IV](https://arxiv.org/html/2609.28553#S4.p1.1)\.
- \[9\]Z\. Xiong, D\. Wang, X\. Liu, F\. Zhong, X\. Wan, X\. Li, Z\. Li, X\. Luo, K\. Chen, H\. Jiang, and M\. Zheng\(2020\)Pushing the boundaries of molecular representation for drug discovery with the graph attention mechanism\.Journal of Medicinal Chemistry63\(16\),pp\. 8749–8760\.External Links:[Document](https://dx.doi.org/10.1021/acs.jmedchem.9b00959)Cited by:[§II](https://arxiv.org/html/2609.28553#S2.p1.1),[TABLE I](https://arxiv.org/html/2609.28553#S4.T1.10.7.1.1),[TABLE III](https://arxiv.org/html/2609.28553#S4.T3.8.1.4.1),[§IV](https://arxiv.org/html/2609.28553#S4.p1.1)\.
- \[10\]W\. Hu, B\. Liu, J\. Gomes, M\. Zitnik, P\. Liang, V\. Pande, and J\. Leskovec\(2020\)Strategies for pre\-training graph neural networks\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§II](https://arxiv.org/html/2609.28553#S2.p1.1),[TABLE I](https://arxiv.org/html/2609.28553#S4.T1.10.3.1.1),[§IV](https://arxiv.org/html/2609.28553#S4.p1.1)\.
- \[11\]P\. Schwaller, T\. Laino, T\. Gaudin, P\. Bolgar, C\. A\. Hunter, C\. Bekas, and A\. A\. Lee\(2019\)Molecular transformer: a model for uncertainty\-calibrated chemical reaction prediction\.ACS Central Science5\(9\),pp\. 1572–1583\.External Links:[Document](https://dx.doi.org/10.1021/acscentsci.9b00576)Cited by:[§II](https://arxiv.org/html/2609.28553#S2.p1.1),[TABLE I](https://arxiv.org/html/2609.28553#S4.T1.10.9.1.1),[§IV](https://arxiv.org/html/2609.28553#S4.p1.1)\.
- \[12\]J\. Ross, B\. Belgodere, V\. Chenthamarakshan, I\. Padhi, Y\. Mroueh, and P\. Das\(2022\)Large\-scale chemical language representations capture molecular structure and properties\.Nature Machine Intelligence4\(12\),pp\. 1256–1264\.External Links:[Document](https://dx.doi.org/10.1038/s42256-022-00580-7)Cited by:[§II](https://arxiv.org/html/2609.28553#S2.p1.1)\.
- \[13\]R\. Zhang, H\. Wen, Z\. Lin, B\. Li, and X\. Zhou\(2025\)Artificial intelligence\-driven drug toxicity prediction: advances, challenges, and future directions\.Toxics13\(7\),pp\. 525\.External Links:[Document](https://dx.doi.org/10.3390/toxics13070525)Cited by:[§II](https://arxiv.org/html/2609.28553#S2.p1.1)\.
- \[14\]H\. Cai, H\. Zhang, D\. Zhao, J\. Wu, and L\. Wang\(2022\)FP\-GNN: a versatile deep learning architecture for enhanced molecular property prediction\.Briefings in Bioinformatics23\(6\),pp\. bbac408\.External Links:[Document](https://dx.doi.org/10.1093/bib/bbac408)Cited by:[§II](https://arxiv.org/html/2609.28553#S2.p2.1)\.
- \[15\]X\. Zang, J\. Zhang, and B\. Tang\(2026\)Molecular representation learning via multimodal fusion and decoupling\.Information Fusion125,pp\. 103493\.External Links:[Document](https://dx.doi.org/10.1016/j.inffus.2025.103493)Cited by:[§II](https://arxiv.org/html/2609.28553#S2.p2.1)\.
- \[16\]G\. W\. Bemis and M\. A\. Murcko\(1996\)The properties of known drugs\. 1\. molecular frameworks\.Journal of Medicinal Chemistry39\(15\),pp\. 2887–2893\.External Links:[Document](https://dx.doi.org/10.1021/jm9602928)Cited by:[§III\-A](https://arxiv.org/html/2609.28553#S3.SS1.p1.1)\.
- \[17\]D\. Weininger\(1988\)SMILES, a chemical language and information system: 1\. introduction to methodology and encoding rules\.Journal of Chemical Information and Computer Sciences28\(1\),pp\. 31–36\.External Links:[Document](https://dx.doi.org/10.1021/ci00057a005)Cited by:[§III\-B1](https://arxiv.org/html/2609.28553#S3.SS2.SSS1.p1.1)\.
- \[18\]K\. Xu, C\. Li, Y\. Tian, T\. Sonobe, K\. Kawarabayashi, and S\. Jegelka\(2018\)Representation learning on graphs with jumping knowledge networks\.InInternational Conference on Machine Learning \(ICML\),pp\. 5453–5462\.Cited by:[§III\-C2](https://arxiv.org/html/2609.28553#S3.SS3.SSS2.p1.1)\.
- \[19\]A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, L\. Kaiser, and I\. Polosukhin\(2017\)Attention is all you need\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§III\-C3](https://arxiv.org/html/2609.28553#S3.SS3.SSS3.p1.1)\.
- \[20\]T\. Lin, P\. Goyal, R\. Girshick, K\. He, and P\. Dollar\(2017\)Focal loss for dense object detection\.InIEEE International Conference on Computer Vision \(ICCV\),pp\. 2980–2988\.Cited by:[§III\-E1](https://arxiv.org/html/2609.28553#S3.SS5.SSS1.p1.1)\.
- \[21\]I\. Loshchilov and F\. Hutter\(2019\)Decoupled weight decay regularization\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§III\-E2](https://arxiv.org/html/2609.28553#S3.SS5.SSS2.p1.1)\.
- \[22\]J\. Davis and M\. Goadrich\(2006\)The relationship between precision\-recall and roc curves\.InProceedings of the 23rd International Conference on Machine Learning \(ICML\),pp\. 233–240\.Cited by:[§IV\-A](https://arxiv.org/html/2609.28553#S4.SS1.p4.1)\.

相似文章

利用基于图的工具改进小型语言模型中的分子性质预测

arXiv cs.AI

本文提出了一种上下文增强提示框架,利用图神经网络专家模型提供预测提示和解释性子图,以改进小型语言模型中的分子性质预测。在MUTAG和Tox21上的实验显示,与仅使用SMILES的基线相比,准确率提升高达74%。