SpecOpt: Contact-Diff Reasoning for Agentic Molecule Optimization Toward Binding Specificity
Summary
SpecOpt introduces a molecular design task to improve drug binding specificity by making targeted modifications to existing compounds using an agentic framework with residue-aware contacts, showing improvement for 84.8% of compounds in a benchmark.
View Cached Full Text
Cached at: 09/21/26, 09:14 AM
# SpecOpt: Contact-Diff Reasoning for Agentic Molecule Optimization Toward Binding Specificity
Source: [https://arxiv.org/html/2609.21165](https://arxiv.org/html/2609.21165)
Thao Nguyen & Heng JiAffiliation:Siebel School of Computing and Data ScienceAffiliation:University of Illinois Urbana\-ChampaignAffiliation:\{thaotn2, hengji\}@illinois\.edu
###### Abstract
Off\-target protein binding is a major source of adverse effects for small\-molecule drugs, yet most structure\-based molecular design methods focus on generating selective compounds*de novo*rather than improving the selectivity of existing, well\-characterized drugs\. We introduce*specificity optimization*\(SpecOpt\), a molecular design task that seeks constrained structural modifications to an existing compound that increase its binding preference for an intended target over known off\-targets while preserving its structural identity and drug\-like properties\. To enable systematic evaluation of this task, we construct a ChEMBL\-derived benchmark from compound–target interaction data, identifying intended targets through curated drug\-mechanism annotations and off\-targets through measured activities\. We then develop an agentic framework for SpecOpt that docks each compound against its intended target and off\-targets, compares the resulting poses through residue\-aware atom–protein contacts, and provides these differential interactions to a large language model to propose targeted structural modifications\. Candidates are retained only if they satisfy molecular similarity, ADMET, and target–off\-target docking selectivity criteria\. On 915 compounds, the agent improves the target–off\-target binding gap for 84\.8% of compounds, shifting the mean gap from−0\.72\-0\.72to\+0\.47\+0\.47kcal/mol while maintaining a mean Tanimoto similarity of0\.720\.72to the starting compounds\. Ablation studies identify residue\-specific contact information as the critical optimization signal: replacing residue identities with binary contact indicators eliminates improvement on all 29/29 ablation compounds\. These results establish SpecOpt as a distinct molecular design problem and demonstrate residue\-aware differential interactions as an effective signal for improving the specificity of existing compounds\.
## 1Introduction
Small\-molecule drugs rarely interact exclusively with their intended targets\. Polypharmacology is widespread\([Keiser et al\., 2007](https://arxiv.org/html/2609.21165#bib.bib14)\), and unintended interactions with off\-target proteins can contribute to adverse effects, dose\-limiting toxicity, and clinical failure\. Improving molecular selectivity is therefore a fundamental objective in drug development\. Recent computational approaches have increasingly addressed selectivity through structure\-based generative design, seeking to generate new molecules that bind strongly to a desired protein while avoiding alternative targets\([Chandra et al\., 2024](https://arxiv.org/html/2609.21165#bib.bib1);[Southiratn et al\., 2025](https://arxiv.org/html/2609.21165#bib.bib3);[Park et al\., 2026](https://arxiv.org/html/2609.21165#bib.bib7)\)\. These methods address an important design problem, but they largely formulate selectivity as a property to optimize during*de novo*molecular generation\.
A different problem arises once a promising drug or drug candidate already exists\. Such a compound may have favorable potency, pharmacokinetic properties, synthetic accessibility, extensive experimental characterization, or even established clinical use, while retaining undesirable activity against one or more off\-targets\([Lounkine et al\., 2012](https://arxiv.org/html/2609.21165#bib.bib20)\)\. In this setting, replacing the molecule with a newly generated compound is not necessarily the desired solution\. Instead, our objective is to modify the existing molecule as little as possible while selectively weakening its off\-target interactions and preserving the properties that made it a viable drug candidate\. This goal builds on evidence that structural modifications can alter off\-target activity\([Wassermann et al\., 2010](https://arxiv.org/html/2609.21165#bib.bib21)\)and on medicinal chemistry principles of addressing liabilities while maintaining on\-target potency and favorable physicochemical properties\([Quancard et al\., 2025](https://arxiv.org/html/2609.21165#bib.bib22)\)\. We refer to this task as*specificity optimization \(SpecOpt\)*: given a starting compound, an intended target, and one or more known off\-targets, identify constrained structural modifications that increase the binding preference for the intended target while preserving the identity and drug\-like characteristics of the starting molecule\.
To our knowledge, specificity optimization of an existing drug has not previously been formulated and systematically studied as a standalone computational molecular design task\. This formulation differs fundamentally from*de novo*selective generation\([Chandra et al\., 2024](https://arxiv.org/html/2609.21165#bib.bib1)\)\. The starting molecule is fixed rather than generated, the permissible structural deviation is constrained, and the optimization objective is inherently comparative: success requires increasing the*specificity gap*between binding to the intended target and binding to off\-target proteins\. Consequently, a successful method must determine not only which molecular structures are compatible with the target, but which modifications to an existing structure can preferentially alter its interactions with one protein over another\.
The fixed starting molecule also makes available a direct structural signal for solving this problem\. The same compound can be docked against its intended target and off\-targets, allowing its binding modes to be compared atom by atom\. Rather than asking only whether a molecular feature contributes to binding, we ask a differential question:*how does each part of the molecule interact with the intended target differently from the off\-targets?*This comparison can identify regions that preferentially support target or off\-target binding and, importantly, regions that occupy similar positions across binding pockets but interact with different residues\. The latter distinction is especially relevant for closely related proteins, where a ligand may contact nearly the same regions of homologous pockets even though the identities and chemical environments of those contacts differ\.
We introduce an agentic framework for this specificity\-optimization task \(Figure[1](https://arxiv.org/html/2609.21165#S1.F1)\)\. Given a compound, its intended target, and known off\-targets, the framework docks the compound against each protein and constructs a residue\-aware, per\-atom representation of the differences between the resulting binding poses\. These signals are aggregated over molecular fragments and provided to a large language model \(LLM\) agent, which reasons over the differential interaction environment to propose targeted structural edits\. The agent can remove fragments that preferentially support off\-target interactions, modify regions exposed to different residue environments, or extend the molecule toward target\-specific interaction opportunities\. Proposed molecules are then evaluated using structural similarity, ADMET, and re\-docking criteria, ensuring that optimization remains anchored to the original compound while improving its predicted specificity\.
Our main contributions are:
- •We formulate specificity optimization as a computational molecular design task\.To our knowledge, this is the first framework to explicitly address the problem of modifying an existing drug or drug candidate to increase its target–off\-target binding preference while constraining structural and drug\-like\-property changes\. This task complements, rather than replaces, existing work on selective*de novo*molecular generation\.
- •We introduce a benchmark and agentic framework for the task\.We construct a ChEMBL\-derived dataset of 1,848 compounds annotated with intended targets and experimentally measured off\-targets and develop an iterative agent combining differential structural analysis, LLM\-guided molecular editing, and property\-based filtering\. On 915 evaluated compounds, the framework improves the docking\-predicted target–off\-target specificity gap for 84\.8% of compounds, while maintaining a mean Tanimoto similarity of 0\.72 to the starting molecules\.
- •We identify residue\-aware differential contacts as the key optimization signal\.Successful optimization depends not merely on whether ligand atoms contact the target or off\-target, but on*which residues*they contact\. Replacing residue\-aware contact differences with a simpler target/off\-target/shared classification eliminates improvement on all 29/29 compounds in our ablation study, highlighting the importance of residue identity for distinguishing interactions across similar binding pockets\.
- •We systematically evaluate additional agentic mechanisms\.We investigate cross\-compound memory, pose\-consistency caution, and block\-synergy mechanisms designed to provide additional experience and structural context\. Although these mechanisms produce their intended intermediate behaviors, they do not measurably improve final optimization outcomes, helping isolate which information is consequential for this task\.
Figure 1:Overview of SpecOpt\.Given a compound, its intended target, and known off\-targets, SpecOpt docks the compound against each protein, computes per\-atom contact differences with residue identities, and uses an LLM agent to propose targeted molecular edits\. Candidate molecules are evaluated using structural similarity, ADMET, and re\-docking criteria, and accepted candidates are iteratively refined in subsequent rounds\.
## 2Related Work
##### Selectivity\-aware molecular generation\.
Recent structure\-based generative methods have incorporated selectivity by optimizing binding to an intended target while penalizing activity against off\-targets\.[Chandra et al\. \(2024\)](https://arxiv.org/html/2609.21165#bib.bib1)use Bayesian optimization over the latent space of a junction\-tree VAE\([Jin et al\., 2018](https://arxiv.org/html/2609.21165#bib.bib12)\), rewarding docking to FLT3 while penalizing docking to homologous kinases\.[Park et al\. \(2026\)](https://arxiv.org/html/2609.21165#bib.bib7)steer molecular generation with dual affinity guidance, while[Southiratn et al\. \(2025\)](https://arxiv.org/html/2609.21165#bib.bib3)use Pareto Monte Carlo tree search over fragment space for the related but inverse objective of dual\-target binding\. These approaches construct molecules*de novo*, whereas SpecOpt begins with an existing compound and seeks constrained structural modifications that improve its binding preference for the intended target while preserving similarity to the original molecule\. This setting also enables direct comparison of how the same molecule interacts with its target and off\-targets\. We compare with[Chandra et al\. \(2024\)](https://arxiv.org/html/2609.21165#bib.bib1)on FLT3 in Section[4\.5](https://arxiv.org/html/2609.21165#S4.SS5), while noting that absolute selectivity is not directly comparable because the two methods address different molecular design settings\.
##### Structure\-based lead optimization\.
Several structure\-based methods modify existing ligands using information from the protein binding pocket\. DeepFrag\([Green et al\., 2021](https://arxiv.org/html/2609.21165#bib.bib17)\)predicts fragment additions from a receptor–ligand complex, while STRIFE\([Hadfield et al\., 2022](https://arxiv.org/html/2609.21165#bib.bib18)\)uses target\-derived pharmacophoric information to guide fragment elaboration\. AutoGrow4\([Spiegel and Durrant, 2020](https://arxiv.org/html/2609.21165#bib.bib19)\)applies docking\-guided genetic search to evolve existing ligands toward improved predicted binding, and FRAME\([Powers et al\., 2023](https://arxiv.org/html/2609.21165#bib.bib2)\)uses geometric deep learning to iteratively grow fragments from a bound ligand, reporting improvements in both predicted affinity and selectivity\. These methods demonstrate that protein structural context can effectively guide modification of an existing molecular starting point\. SpecOpt differs in explicitly reasoning over the*difference*between how the same molecule interacts with its intended target and known off\-targets, and using these differential interactions to guide constrained molecular edits\. To our knowledge, SpecOpt is the first framework to formulate specificity optimization of an existing compound against known off\-targets as a standalone molecular design task\.
##### LLM agents for molecular design\.
Tool\-augmented LLMs have been applied broadly to chemistry\([M\. Bran et al\., 2024](https://arxiv.org/html/2609.21165#bib.bib13)\)and more recently to lead optimization\.[Li et al\. \(2026\)](https://arxiv.org/html/2609.21165#bib.bib4)formulate tool selection as sequential decision\-making over molecular edit trajectories, while[Wang et al\. \(2026\)](https://arxiv.org/html/2609.21165#bib.bib5)incorporate memory into an agentic reinforcement learning optimizer\. MolLingo\([Nguyen and Ji, 2026](https://arxiv.org/html/2609.21165#bib.bib6)\)further demonstrates that LLMs can reason over chemically meaningful molecular fragments together with docking\-derived residue and binding\-pocket context to propose molecular modifications that improve predicted target binding\. Building on this capability, SpecOpt focuses specifically on target\-versus\-off\-target specificity, providing the LLM with differential structural information between target\- and off\-target\-bound poses to guide molecular editing\.
## 3SpecOpt: Contact\-Diff\-Guided Specificity Optimization
### 3\.1Task Formulation
Given a compoundmm, an intended targetTT, and a set of off\-targets𝒪\\mathcal\{O\}, lets\(m,P\)s\(m,P\)be the AutoDock Vina\([Trott and Olson, 2010](https://arxiv.org/html/2609.21165#bib.bib8)\)score ofmmagainst proteinPP\(more negative = tighter binding\)\. We define thespecificity gap
g\(m\)=minO∈𝒪s\(m,O\)−s\(m,T\),g\(m\)\\;=\\;\\min\_\{O\\in\\mathcal\{O\}\}s\(m,O\)\\;\-\\;s\(m,T\),\(1\)the margin between the target and the*tightest\-binding*off\-target\. A positive gap means the molecule prefers its target over every off\-target; the optimization objective is to increaseggrelative to the original compoundm0m\_\{0\}, subject to staying chemically close tom0m\_\{0\}\.
### 3\.2Benchmark construction
We derive a specificity benchmark from ChEMBL\([Mendez et al\., 2019](https://arxiv.org/html/2609.21165#bib.bib9)\)’s compound–target interaction data\. A compound’s*intended*targets are those with a curated drug\-mechanism designation; its*off\-targets*are proteins with measured activity but no such designation\. Constructing this cleanly required correcting three systematic over\-counting problems: mutation panels, where different point mutations of the same protein are recorded as separate targets but are collapsed here because we focus on cross\-protein specificity rather than selectivity among protein variants; protein complexes and subunits, where a subunit is counted separately from its parent complex; and generic\-organism duplicates, where the same protein is registered under both a specific and an unspecified organism\. The resulting dataset contains1,848 compounds, with a median of 3 and a mean of 7\.5 off\-targets each\.
For docking\-based experiments we filter to compounds where both the intended target and at least one off\-target are single proteins with a resolvable experimental or predicted structure, and additionally exclude conformationally extreme molecules \(\>45\>45heavy atoms or\>10\>10rotatable bonds\), since docking cost scales with ligand flexibility\. This yields a working pool of916 compounds\. To make repeated experimentation tractable we precompute baseline docking for all 15,195 dockable \(compound, target\) pairs in the dataset, of which13,636\(89\.7%\) resolve successfully\.
### 3\.3Differential interaction analysis
The central signal in SpecOpt comes from comparing how the same ligand interacts with its intended target and off\-targets\. For each ligand atomii, we identify protein residues within 7 Å in the target\-bound pose and each off\-target\-bound pose\. Based on these contacts, atoms are classified astarget\-only,offtarget\-only,shared\-same\-residues, orshared\-different\-residues\. The last two categories distinguish atoms that contact the same residue identities across proteins from those that occupy similar regions of the binding pockets but interact with different residues\. Independently, an atom is marked as agrowth\-opportunitywhen there is measurably more unoccupied pocket volume near that atom in the target than in the off\-target\.
We aggregate these atom\-level signals over BRICS\([Degen et al\., 2008](https://arxiv.org/html/2609.21165#bib.bib10)\)fragments and assign each fragment an editing suggestion:remove,keep,redesign,grow, orneutral\. These suggestions determine how each fragment and its local interaction environment are presented to GPT\-5\.4\([OpenAI, 2026](https://arxiv.org/html/2609.21165#bib.bib15)\), which serves as our LLM agent\. Importantly, contact with both proteins does not imply that an atom is uninformative: when the contacted residues differ between the target and off\-target, the position provides an opportunity for a chemical modification to preferentially affect one interaction environment\. As shown in Section[4\.3](https://arxiv.org/html/2609.21165#S4.SS3), preserving this residue\-level distinction is critical to the effectiveness of SpecOpt\.
Figure 2:Differential interaction analysis and molecular editing in SpecOpt\.Target\- and off\-target\-bound poses provide atom\-level contact differences and growth opportunities, which are aggregated into fragment\-level editing suggestions to guide candidate molecule construction\.
### 3\.4Iterative optimization and candidate filtering
Starting from the original compound, SpecOpt performs iterative optimization by alternating between LLM\-guided molecular editing and candidate evaluation\. At each round, the fragment\-level interaction summary, current docking scores, and previously attempted edits are provided to the LLM, which proposes fragment\-level modifications\. These proposals are applied to generate candidate molecules, with one fragment modified at a time\. Each candidate must pass three criteria before acceptance: Tanimoto similarity to the*original*molecule above 0\.4, ADMET changes within tolerance for direction\-unambiguous toxicity endpoints \(AMES, DILI, hERG, and P\-gp\), and an improved target–off\-target specificity gap after re\-docking\. We evaluate ADMET properties using the prediction model from mCLM\([Edwards et al\., 2026](https://arxiv.org/html/2609.21165#bib.bib16)\)\. Accepted candidates become the starting molecules for the next round, allowing structural modifications to accumulate over successive iterations\. We additionally evaluate two extensions in Section[4\.3](https://arxiv.org/html/2609.21165#S4.SS3): a*combo*mechanism that combines independently beneficial edits, and a two\-level memory mechanism consisting of within\-compound edit history and cross\-compound memory retrieval\.
Because both docking and LLM generation introduce stochasticity, we fix the RDKit\([Landrum, n\.d\.](https://arxiv.org/html/2609.21165#bib.bib11)\)conformer seed and Vina search seed throughout the experiments\. Fixing both is necessary for reproducible docking scores; otherwise, identical inputs can produce slightly different scores \(e\.g\.,−7\.1\-7\.1vs\.−7\.2\-7\.2kcal/mol\), which can in turn change which edits are selected\. LLM proposals retain residual stochasticity, which defines the noise floor against which the ablation results are interpreted\.
## 4Experiments
### 4\.1Main result
We run the full system over the 916\-compound pool with a budget of 8 candidate proposals per round and 4 rounds per compound\. Failed runs are retried up to two additional times, for a maximum of three attempts per compound; a compound is excluded if all three attempts fail\. Across 945 attempts, including retries, 915 distinct compounds complete successfully, while one compound exhausts the retry limit\. The 30 failed attempts \(3\.2% of all attempts\) are almost all due to baseline\-docking failures on large, flexible molecules\. All subsequent main results are computed over the 915 successfully completed compounds\.
Table 1:Main result over the full compound pool\. Gap is defined in Eq\.[1](https://arxiv.org/html/2609.21165#S3.E1); positive means target\-preferring\. The paired test compares each compound’s own baseline against its own optimized result\.Table[1](https://arxiv.org/html/2609.21165#S4.T1)and Figure[3](https://arxiv.org/html/2609.21165#S4.F3)summarize the overall results\. On average, compounds shift from preferential binding to off\-targets at baseline to preferential binding to their intended targets after optimization\. Among the 575 compounds with a negative baseline specificity gap, 190 \(33%\) achieve a positive gap after optimization\. The similarity distribution \(Figure[3](https://arxiv.org/html/2609.21165#S4.F3)c\) shows that these improvements are achieved while largely preserving the original molecular structures\.
Figure 3:Main result across 915 compounds\.\(a\) The specificity\-gap distribution shifts from predominantly negative \(off\-target\-preferring\) to predominantly positive\. \(b\) Per\-compound improvement\. \(c\) Structural conservation; the spike at1\.01\.0is compounds returned unchanged because no candidate passed the admission gates\.
### 4\.2Optimization examples and case studies
Because SpecOpt records the contact\-difference signal, LLM proposals, and accepted edits at each round, individual optimization trajectories are fully traceable\. Figure[4](https://arxiv.org/html/2609.21165#S4.F4)illustrates this process for edoxaban, a factor Xa inhibitor with prothrombin as an off\-target\. The contact\-difference analysis identifies the cyclohexane linker as agrowregion, with three atoms having greater available pocket volume in factor Xa, while two other fragments are marked forredesign\. Consistent with this signal, the LLM proposes extending the cyclohexane ring\. A methyl and fluorine modification are accepted in the first round, followed by addition of a trifluoromethyl group at the same position in the second round\. Across these edits, the specificity gap improves from−5\.6\-5\.6to\+1\.5\+1\.5kcal/mol while retaining a Tanimoto similarity of0\.760\.76to the original compound\.
Contact\-difference summary supplied to the LLM \(edoxaban, factor Xa vs\. prothrombin\):•\[1\*\]\[C@@H\]1C\[C@@H\]\(C\(=O\)N\(C\)C\)CC\[C@@H\]1\[2\*\]: contacts are shared with the off\-target, but3 atoms have substantially more available space in the target pocket; candidate togrow\.•\[2\*\]c1nc2c\(s1\)CN\(C\)CC2: contacts occur in both poses at 7 atoms, butagainst different residues; candidate toredesign\.•\[1\*\]C\(=O\)N\[2\*\]: contacts occur in both poses at 3 atoms, against different residues; candidate toredesign\.•\[1\*\]Nc1ccc\(Cl\)cn1,\[1\*\]NC\(=O\)C\(\[2\*\]\)=O: no docking contact data available\.Accepted edits:•Round 1: extend cyclohexane with methyl and fluorine modifications \(gap−5\.6→−1\.0\-5\.6\\rightarrow\-1\.0\)\.•Round 2: add CF3at the same position \(gap−1\.0→\+1\.5\-1\.0\\rightarrow\+1\.5\)\.Final:gap−5\.6→\+1\.5\-5\.6\\rightarrow\+1\.5kcal/mol; Tanimoto similarity0\.760\.76\.
Figure 4:Example optimization trajectory for edoxaban\. The contact\-difference analysis identifies candidate regions for modification, and the LLM proposes corresponding structural edits that are evaluated and accumulated across optimization rounds\.Table[2](https://arxiv.org/html/2609.21165#S4.T2)presents additional examples that begin with negative specificity gaps, achieve positive gaps after optimization, and retain Tanimoto similarity≥0\.6\\geq 0\.6to the original compound\. The examples span multiple target classes, including proteases, GPCRs, kinases, and metabolic enzymes\. The resulting modifications are generally small, localized substitutions, including fluorine and CF3additions, nitrile addition, substituent repositioning, and methylation, consistent with the constrained lead\-optimization setting of SpecOpt\.
Figure 5:Selected SpecOpt case studies showing the original \(left\) and optimized \(right\) structures\. Highlighted atoms indicate structural changes relative to the maximum common substructure of the original compound\. Specificity gaps are reported in kcal/mol\.Table 2:Selected case studies from the main evaluation\. “Edit” summarizes the net structural change relative to the original compound\.
### 4\.3Ablation study: identifying the key signal
To isolate the contribution of each component, we fix a 30\-compound subset stratified by baseline specificity gap and evaluate the full system against five ablated variants\. Each ablation is compared with the full\-system result on the same compound \(Figure[6](https://arxiv.org/html/2609.21165#S4.F6)a\)\.
Figure 6:Ablation and signal analysis\. \(a\) Mean change in specificity\-gap improvement when each component is removed; positive values indicate that the component improves performance\. Green indicatesp<0\.05p<0\.05, orangep<0\.1p<0\.1, and grey no statistical significance\. \(b\) Fragment\-level interaction signals aggregated across the 915\-compound evaluation\. Exclusive target\- or off\-target\-contact signals do not occur, while most informative signals arise from different residue identities across shared contacts and target\-specific growth opportunities\.##### Residue identity provides the critical optimization signal\.
Replacing the residue\-aware classification with a naive three\-way classification \(target\-only, off\-target\-only, or shared contact\) reduces the mean specificity\-gap improvement from\+0\.903\+0\.903to0\.0000\.000across all 29/29 compounds: without residue identity, the system detects no differentiating interaction signal and terminates without modification\. Figure[6](https://arxiv.org/html/2609.21165#S4.F6)b explains this behavior\. Across the full 915\-compound evaluation, the dominant signals areshared\-different\-residues\(5,668 fragment classifications\) andgrowth\-opportunity\(2,321\)\. Thus, simply identifying whether a ligand region contacts the target or off\-target is insufficient; the useful signal comes primarily from differences in*which residues*are contacted\.
This observation also explains why the deterministicremoveoperation contributes no edits\. The operation is designed to delete fragments that interact exclusively with an off\-target, but its trigger condition does not occur in our evaluation\. We therefore retain this operation in the framework but report its lack of activation rather than relaxing its definition post hoc\.
Table 3:Performance grouped by the relationship between the intended target and the off\-target defining the specificity gap\. Performance is consistent across relationship types\.
##### Mechanisms with no measurable performance benefit\.
We evaluate three additional mechanisms that operate as intended but do not significantly improve optimization performance\. \(i\)*Cross\-compound memory*retrieves an average of 2\.5 relevant records from previous compounds, but removing it has no significant effect on performance \(\+0\.052\+0\.052,p=0\.65p=0\.65\)\. To test whether this is caused by limited relevance of the retrieved information, we repeat the experiment on 24 drugs with identical target and off\-target sets, where cross\-compound information should be maximally transferable, but again observe no benefit \(−0\.067\-0\.067,p=0\.33p=0\.33\)\. \(ii\) The*pose\-consistency*mechanism identifies cases where a ligand adopts substantially different conformations in the target and off\-target pockets and warns the LLM accordingly\. We evaluate it on flexible molecules, where 10/17 compounds receive medium\- or high\-risk warnings, but observe no improvement \(−0\.07\-0\.07\)\. \(iii\) The*combo*mechanism attempts to combine individually beneficial edits and activates for approximately 60% of eligible compounds\. Most combined edits \(6/8\) fail the ADMET acceptance criteria, and the overall improvement is small and not statistically significant \(\+0\.079\+0\.079,p=0\.08p=0\.08\)\.
##### A budget\-dependent effect\.
Within\-compound memory, which prevents the agent from repeating previously attempted edits, shows no measurable benefit under a 4\-round optimization budget \(−0\.111\-0\.111,p=0\.97p=0\.97\)\. However, when the budget is increased to 8 rounds, memory produces a larger improvement \(\+0\.603\+0\.603,p=0\.091p=0\.091\)\. The memory\-enabled condition also evaluates more candidates on average \(27\.327\.3vs\.23\.623\.6\), suggesting that avoiding repeated proposals allows the optimization to continue exploring rather than triggering the “no improvement” early\-stopping criterion\. Because the main evaluation uses the 4\-round budget, it may underestimate the benefit of within\-compound memory for longer optimization runs\.
### 4\.4Multiple off\-targets
For compounds with multiple off\-targets, we retain up to three off\-targets based first on structural availability and then on binding affinity\. The contact\-difference signals from these off\-targets are pooled to construct a single fragment\-level summary for the LLM, allowing SpecOpt to optimize against multiple competing proteins simultaneously\. To examine the effect on a broader set of off\-targets, we additionally re\-dock the FLT3 benchmark compounds against five common off\-targets \(Appendix[A\.1](https://arxiv.org/html/2609.21165#A1.SS1), Figure[7](https://arxiv.org/html/2609.21165#A1.F7)\)\. On average, optimization strengthens binding to the intended target while weakening binding to most off\-targets\. Only5/205/20compounds weaken all five off\-targets simultaneously, while others show mixed effects\. This behavior is expected because the current implementation optimizes against at most three off\-targets; proteins outside this selected set are not explicitly considered during optimization\. Extending SpecOpt to jointly optimize against a larger off\-target panel is a natural direction for future work\.
### 4\.5External comparison: FLT3
To compare SpecOpt with prior selectivity\-aware molecular design, we adopt the setting of[Chandra et al\. \(2024\)](https://arxiv.org/html/2609.21165#bib.bib1), using FLT3 as the intended target and PDGFRA, KIT, VEGFR2, MK2, and JAK2 as off\-targets\. Rather than optimizing newly generated molecules, we apply SpecOpt to 24 real FLT3\-targeting drugs in our benchmark, including sunitinib, midostaurin, gilteritinib, and crenolanib\. Among 23 successful runs, SpecOpt improves the specificity gap for 20 compounds, shifting the mean gap from−2\.03\-2\.03to−0\.91\-0\.91kcal/mol \(p=8\.8×10−5p=8\.8\\times 10^\{\-5\}\)\.
Absolute selectivity values are not directly comparable between the two methods because they address different tasks and use different docking pipelines\.[Chandra et al\. \(2024\)](https://arxiv.org/html/2609.21165#bib.bib1)generate molecules*de novo*, whereas SpecOpt modifies existing drugs under structural constraints\.
## 5Discussion and Limitations
Our results highlight the importance of residue\-aware differential contacts rather than contact presence alone for constrained molecular optimization\. The LLM translates these signals into specific chemical edits, and richer descriptions of the local residue environment may further improve this process\. Additional agentic mechanisms do not consistently improve performance, emphasizing the need to validate their contributions rather than assume that greater complexity is beneficial\. Several limitations remain\. First, SpecOpt optimizes against at most three selected off\-targets\. Evaluation against five off\-targets shows that improvements can extend beyond this set, but some compounds also strengthen binding to individual off\-targets; improved specificity against the selected proteins therefore does not guarantee broad selectivity\. Second, all reported binding improvements are based on Vina docking scores\. Fixed docking seeds support reproducible evaluation but do not establish experimental affinity or selectivity, which require experimental validation or assessment with higher\-accuracy free\-energy calculations\.
## 6Conclusion
We introduced*specificity optimization*\(SpecOpt\), the task of modifying an existing drug to improve its binding preference for an intended target over known off\-targets while preserving the original molecule\. We constructed a benchmark for this task and developed an agentic framework combining differential structural analysis with LLM\-guided molecular editing\. Our results identify residue\-aware differential contacts as the key optimization signal and demonstrate that existing drugs can be systematically modified toward improved predicted specificity while retaining substantial structural similarity\. SpecOpt therefore provides a complementary direction to selective*de novo*generation and a foundation for further development of specificity\-aware lead optimization\.
#### Reproducibility Statement
The benchmark construction, filtering criteria, agent implementation, ablation toggles, and the exact per\-compound outputs behind every number in this paper are contained in the accompanying code repository\. All docking uses fixed seeds; the precomputed docking cache is released to make the main run reproducible without repeating∼\\sim15k docking calls\. Code is available at[https://anonymous\.4open\.science/r/specopt\-611E](https://anonymous.4open.science/r/specopt-611E)\.
#### Acknowledgments
This work was supported by the NSF Molecule Maker Lab Institute \(MMLI\), an AI Institute for Molecular Discovery, Synthesis Strategy, and Manufacturing, funded by the U\.S\. National Science Foundation under Awards No\. 2019897 and 2505932\.
## References
- H\. M\. Berman, J\. Westbrook, Z\. Feng, G\. Gilliland, T\. N\. Bhat, H\. Weissig, I\. N\. Shindyalov, and P\. E\. BourneThe protein data bank\.Nucleic Acids Research28\(1\),pp\. 235–242\.External Links:[Document](https://dx.doi.org/10.1093/nar/28.1.235)Cited by:[§A\.2](https://arxiv.org/html/2609.21165#A1.SS2.SSS0.Px2.p1.1),[§A\.3](https://arxiv.org/html/2609.21165#A1.SS3.SSS0.Px1.p1.1)\.
- Bickertonet al\.\(2012\)G\. R\. Bickerton, G\. V\. Paolini, J\. Besnard, S\. Muresan, and A\. L\. HopkinsQuantifying the chemical beauty of drugs\.Nature Chemistry4,pp\. 90–98\.External Links:[Document](https://dx.doi.org/10.1038/nchem.1243)Cited by:[§A\.2](https://arxiv.org/html/2609.21165#A1.SS2.SSS0.Px1.p2.1)\.
- Chandraet al\.\(2024\)R\. Chandra, R\. I\. Horne, and M\. VendruscoloBayesian optimization in the latent space of a variational autoencoder for the generation of selective FLT3 inhibitors\.Journal of Chemical Theory and Computation20\(1\),pp\. 469–476\.External Links:[Document](https://dx.doi.org/10.1021/acs.jctc.3c01224)Cited by:[§1](https://arxiv.org/html/2609.21165#S1.p1.1),[§1](https://arxiv.org/html/2609.21165#S1.p3.1),[§2](https://arxiv.org/html/2609.21165#S2.SS0.SSS0.Px1.p1.1),[§4\.5](https://arxiv.org/html/2609.21165#S4.SS5.p1.1),[§4\.5](https://arxiv.org/html/2609.21165#S4.SS5.p2.1)\.
- Degenet al\.\(2008\)J\. Degen, C\. Wegscheid\-Gerlach, A\. Zaliani, and M\. RareyOn the art of compiling and using ’drug\-like’ chemical fragment spaces\.ChemMedChem3\(10\),pp\. 1503–1507\.External Links:[Document](https://dx.doi.org/10.1002/cmdc.200800178)Cited by:[§A\.6](https://arxiv.org/html/2609.21165#A1.SS6.p1.1),[§3\.3](https://arxiv.org/html/2609.21165#S3.SS3.p2.1)\.
- Edwardset al\.\(2026\)C\. Edwards, C\. Han, G\. Lee, T\. Nguyen, S\. Szymkuć, C\. K\. Prasad, B\. Jin, J\. Han, Y\. Diao, G\. Liu, H\. Peng, B\. A\. Grzybowski, M\. D\. Burke, and H\. JimCLM: a modular chemical language model that generates functional and makeable molecules\.InInternational Conference on Learning Representations \(ICLR\),External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2026/hash/76f2614f3cabc22b8a1e0d9a6aa35b61-Abstract-Conference.html)Cited by:[§3\.4](https://arxiv.org/html/2609.21165#S3.SS4.p1.1)\.
- Greenet al\.\(2021\)H\. Green, D\. R\. Koes, and J\. D\. DurrantDeepFrag: a deep convolutional neural network for fragment\-based lead optimization\.Chemical Science12\(23\),pp\. 8036–8047\.External Links:[Document](https://dx.doi.org/10.1039/D1SC00163A)Cited by:[§2](https://arxiv.org/html/2609.21165#S2.SS0.SSS0.Px2.p1.1)\.
- Hadfieldet al\.\(2022\)T\. E\. Hadfield, F\. Imrie, A\. Merritt, K\. Birchall, and C\. M\. DeaneIncorporating target\-specific pharmacophoric information into deep generative models for fragment elaboration\.Journal of Chemical Information and Modeling62\(10\),pp\. 2280–2292\.External Links:[Document](https://dx.doi.org/10.1021/acs.jcim.1c01311)Cited by:[§2](https://arxiv.org/html/2609.21165#S2.SS0.SSS0.Px2.p1.1)\.
- Halgren \(1996\)T\. A\. HalgrenMerck molecular force field\. I\. basis, form, scope, parameterization, and performance of MMFF94\.Journal of Computational Chemistry17\(5–6\),pp\. 490–519\.External Links:[Document](https://dx.doi.org/10.1002/%28SICI%291096-987X%28199604%2917%3A5/6%3C490%3A%3AAID-JCC1%3E3.0.CO%3B2-P)Cited by:[§A\.3](https://arxiv.org/html/2609.21165#A1.SS3.SSS0.Px3.p1.1)\.
- Jinet al\.\(2018\)W\. Jin, R\. Barzilay, and T\. JaakkolaJunction tree variational autoencoder for molecular graph generation\.Proceedings of the 35th International Conference on Machine Learning \(ICML\)\.Cited by:[§2](https://arxiv.org/html/2609.21165#S2.SS0.SSS0.Px1.p1.1)\.
- Jumperet al\.\(2021\)J\. Jumper, R\. Evans, A\. Pritzel,et al\.Highly accurate protein structure prediction with AlphaFold\.Nature596,pp\. 583–589\.External Links:[Document](https://dx.doi.org/10.1038/s41586-021-03819-2)Cited by:[§A\.2](https://arxiv.org/html/2609.21165#A1.SS2.SSS0.Px2.p1.1)\.
- Keiseret al\.\(2007\)M\. J\. Keiser, B\. L\. Roth, B\. N\. Armbruster, P\. Ernsberger, J\. J\. Irwin, and B\. K\. ShoichetRelating protein pharmacology by ligand chemistry\.Nature Biotechnology25\(2\),pp\. 197–206\.External Links:[Document](https://dx.doi.org/10.1038/nbt1284)Cited by:[§1](https://arxiv.org/html/2609.21165#S1.p1.1)\.
- Landrum \(n\.d\.\)G\. LandrumRDKit: open\-source cheminformatics\.Note:[https://www\.rdkit\.org](https://www.rdkit.org/)Accessed 2026Cited by:[§A\.6](https://arxiv.org/html/2609.21165#A1.SS6.p1.1),[§3\.4](https://arxiv.org/html/2609.21165#S3.SS4.p2.1)\.
- Le Guillouxet al\.\(2009\)V\. Le Guilloux, P\. Schmidtke, and P\. TufferyFpocket: an open source platform for ligand pocket detection\.BMC Bioinformatics10,pp\. 168\.External Links:[Document](https://dx.doi.org/10.1186/1471-2105-10-168)Cited by:[item 2](https://arxiv.org/html/2609.21165#A1.I1.i2.p1.1)\.
- Liet al\.\(2026\)L\. Li, H\. Zhang, R\. Fan, B\. Chen, and J\. ZhouMolecular lead optimization via agentic tool planning\.arXiv preprint arXiv:2605\.28862\.Cited by:[§2](https://arxiv.org/html/2609.21165#S2.SS0.SSS0.Px3.p1.1)\.
- Lounkineet al\.\(2012\)E\. Lounkine, M\. J\. Keiser, S\. Whitebread, D\. Mikhailov, J\. Hamon, J\. L\. Jenkins, P\. Lavan, E\. Weber, A\. K\. Doak, S\. Côté, B\. K\. Shoichet, and L\. UrbanLarge\-scale prediction and testing of drug activity on side\-effect targets\.Nature486\(7403\),pp\. 361–367\.External Links:[Document](https://dx.doi.org/10.1038/nature11159)Cited by:[§1](https://arxiv.org/html/2609.21165#S1.p2.1)\.
- M\. Branet al\.\(2024\)A\. M\. Bran, S\. Cox, O\. Schilter, C\. Baldassari, A\. D\. White, and P\. SchwallerAugmenting large language models with chemistry tools\.Nature Machine Intelligence6,pp\. 525–535\.External Links:[Document](https://dx.doi.org/10.1038/s42256-024-00832-8)Cited by:[§2](https://arxiv.org/html/2609.21165#S2.SS0.SSS0.Px3.p1.1)\.
- Mendezet al\.\(2019\)D\. Mendez, A\. Gaulton, A\. P\. Bento, J\. Chambers, M\. De Veij, E\. Félix, M\. P\. Magariños, J\. F\. Mosquera, P\. Mutowo, M\. Nowotka,et al\.ChEMBL: towards direct deposition of bioassay data\.Nucleic Acids Research47\(D1\),pp\. D930–D940\.External Links:[Document](https://dx.doi.org/10.1093/nar/gky1075)Cited by:[§3\.2](https://arxiv.org/html/2609.21165#S3.SS2.p1.1)\.
- Morriset al\.\(2009\)G\. M\. Morris, R\. Huey, W\. Lindstrom, M\. F\. Sanner, R\. K\. Belew, D\. S\. Goodsell, and A\. J\. OlsonAutoDock4 and AutoDockTools4: automated docking with selective receptor flexibility\.Journal of Computational Chemistry30\(16\),pp\. 2785–2791\.External Links:[Document](https://dx.doi.org/10.1002/jcc.21256)Cited by:[§A\.3](https://arxiv.org/html/2609.21165#A1.SS3.p1.1)\.
- Nguyen and Ji \(2026\)T\. Nguyen and H\. JiMolLingo: molecule\-native representations for LLM\-powered scientific agents\.arXiv preprint arXiv:2605\.27853\.Cited by:[§2](https://arxiv.org/html/2609.21165#S2.SS0.SSS0.Px3.p1.1)\.
- OpenAI \(2026\)OpenAIIntroducing GPT\-5\.4\.Note:[https://openai\.com/index/introducing\-gpt\-5\-4/](https://openai.com/index/introducing-gpt-5-4/)Cited by:[§3\.3](https://arxiv.org/html/2609.21165#S3.SS3.p2.1)\.
- Parket al\.\(2026\)H\. Park, H\. Kim, S\. Choi, S\. Lee, Y\. Kim, and S\. ParkTheSelective: dual affinity\-guided diffusion for selective molecular generation\.InAdvances in Knowledge Discovery and Data Mining \(PAKDD 2026\),Lecture Notes in Computer Science, Vol\.16598,pp\. 16–27\.External Links:[Document](https://dx.doi.org/10.1007/978-981-92-1462-4%5F2)Cited by:[§1](https://arxiv.org/html/2609.21165#S1.p1.1),[§2](https://arxiv.org/html/2609.21165#S2.SS0.SSS0.Px1.p1.1)\.
- Powerset al\.\(2023\)A\. S\. Powers, H\. H\. Yu, P\. Suriana, R\. V\. Koodli, T\. Lu, J\. M\. Paggi, and R\. O\. DrorGeometric deep learning for structure\-based ligand design\.ACS Central Science9\(12\),pp\. 2257–2267\.External Links:[Document](https://dx.doi.org/10.1021/acscentsci.3c00572)Cited by:[§2](https://arxiv.org/html/2609.21165#S2.SS0.SSS0.Px2.p1.1)\.
- Quancardet al\.\(2025\)J\. Quancard, A\. Bach, C\. Borsari, R\. Craft, C\. Gnamm, S\. M\. Guéret, I\. V\. Hartung, H\. F\. Koolman, S\. Laufer, S\. Lepri, J\. Messinger, K\. Ritter, G\. Sbardella, A\. Unzue Lopez, M\. K\. Willwacher, B\. Cox, and R\. J\. YoungThe european federation for medicinal chemistry and chemical biology \(EFMC\) best practice initiative: hit to lead\.ChemMedChem20\(8\),pp\. e202400931\.External Links:[Document](https://dx.doi.org/10.1002/cmdc.202400931)Cited by:[§1](https://arxiv.org/html/2609.21165#S1.p2.1)\.
- Schmidtke and Barril \(2010\)P\. Schmidtke and X\. BarrilUnderstanding and predicting druggability\. a high\-throughput method for detection of drug binding sites\.Journal of Medicinal Chemistry53\(15\),pp\. 5858–5867\.External Links:[Document](https://dx.doi.org/10.1021/jm100574m)Cited by:[item 2](https://arxiv.org/html/2609.21165#A1.I1.i2.p1.1)\.
- Southiratnet al\.\(2025\)T\. Southiratn, B\. Koo, Y\. Lu, and S\. KimCombiMOTS: combinatorial multi\-objective tree search for dual\-target molecule generation\.InProceedings of the 42nd International Conference on Machine Learning \(ICML\),Proceedings of Machine Learning Research, Vol\.267,pp\. 56650–56691\.Cited by:[§1](https://arxiv.org/html/2609.21165#S1.p1.1),[§2](https://arxiv.org/html/2609.21165#S2.SS0.SSS0.Px1.p1.1)\.
- Spiegel and Durrant \(2020\)J\. O\. Spiegel and J\. D\. DurrantAutoGrow4: an open\-source genetic algorithm for de novo drug design and lead optimization\.Journal of Cheminformatics12,pp\. 25\.External Links:[Document](https://dx.doi.org/10.1186/s13321-020-00429-4)Cited by:[§2](https://arxiv.org/html/2609.21165#S2.SS0.SSS0.Px2.p1.1)\.
- The UniProt Consortium \(2025\)The UniProt ConsortiumUniProt: the universal protein knowledgebase in 2025\.Nucleic Acids Research53\(D1\),pp\. D609–D617\.External Links:[Document](https://dx.doi.org/10.1093/nar/gkae1010)Cited by:[§A\.2](https://arxiv.org/html/2609.21165#A1.SS2.SSS0.Px2.p1.1)\.
- Trott and Olson \(2010\)O\. Trott and A\. J\. OlsonAutoDock Vina: improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading\.Journal of Computational Chemistry31\(2\),pp\. 455–461\.External Links:[Document](https://dx.doi.org/10.1002/jcc.21334)Cited by:[§3\.1](https://arxiv.org/html/2609.21165#S3.SS1.p1.1)\.
- Wanget al\.\(2020\)S\. Wang, J\. Witek, G\. A\. Landrum, and S\. RinikerImproving conformer generation for small rings and macrocycles based on distance geometry and experimental torsional\-angle preferences\.Journal of Chemical Information and Modeling60,pp\. 2044–2058\.External Links:[Document](https://dx.doi.org/10.1021/acs.jcim.0c00025)Cited by:[§A\.2](https://arxiv.org/html/2609.21165#A1.SS2.SSS0.Px2.p1.1),[§A\.3](https://arxiv.org/html/2609.21165#A1.SS3.SSS0.Px3.p1.1)\.
- Wanget al\.\(2026\)Z\. Wang, Y\. Wen, A\. Pandey, H\. Liu, and K\. DingMolMem: memory\-augmented agentic reinforcement learning for sample\-efficient molecular optimization\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),San Diego, California, United States,pp\. 43694–43712\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.acl-long.2024),[Link](https://aclanthology.org/2026.acl-long.2024/)Cited by:[§2](https://arxiv.org/html/2609.21165#S2.SS0.SSS0.Px3.p1.1)\.
- Wassermannet al\.\(2010\)A\. M\. Wassermann, L\. Peltason, and J\. BajorathComputational analysis of multi\-target structure–activity relationships to derive preference orders for chemical modifications toward target selectivity\.ChemMedChem5\(6\),pp\. 847–858\.External Links:[Document](https://dx.doi.org/10.1002/cmdc.201000064)Cited by:[§1](https://arxiv.org/html/2609.21165#S1.p2.1)\.
## Appendix AAppendix
### A\.1Generalization across a broader off\-target panel
Figure[7](https://arxiv.org/html/2609.21165#A1.F7)examines whether optimization generalizes across a broader off\-target panel\. Most compounds show weaker predicted binding to multiple off\-targets, and five compounds weaken all five simultaneously, indicating that improvements can generalize beyond individual proteins\. However, several compounds strengthen binding to at least one off\-target\. This mixed behavior is expected because SpecOpt currently considers at most three off\-targets during optimization, whereas the figure evaluates five; off\-targets outside the selected optimization set are not explicitly constrained\.
Figure 7:Binding\-score changes after optimization for FLT3\-targeting compounds\. Left: change in target binding, where negative values indicate stronger binding\. Right: changes across five off\-targets, where positive values \(blue\) indicate weaker binding\. Five compounds weaken all five off\-targets simultaneously, while the remaining compounds show mixed effects across off\-targets\.
### A\.2Benchmark characterization and docking validation
We characterize the benchmark composition and assess the reliability and limitations of its docking scores\.
##### Composition\.
The benchmark contains 1,848 compounds spanning 994 distinct protein targets \(Figure[8](https://arxiv.org/html/2609.21165#A1.F8)\)\. The number of annotated off\-targets per compound follows a heavy\-tailed distribution, with a median of 3, a mean of 7\.5, a 90th percentile of 18, and a maximum of 148\. This variation is relevant to optimization because the specificity gap is defined by the tightest\-binding off\-target\. Compounds with larger off\-target sets may present more competing interactions, increasing the possibility that reducing binding to one off\-target leaves another as the principal determinant of the specificity gap \(Section[4\.4](https://arxiv.org/html/2609.21165#S4.SS4)\)\.
The benchmark consists of drugs and drug candidates rather than generated molecules: 1,214 of 1,848 compounds \(66%\) are approved \(ChEMBLmax\_phase4\), with the remainder in Phase 1–3\. The median molecular weight is 383 Da, and the median quantitative estimate of drug\-likeness \(QED\)\([Bickerton et al\., 2012](https://arxiv.org/html/2609.21165#bib.bib23)\)is 0\.57\. These baseline properties motivate the use of similarity and ADMET constraints to preserve characteristics of the starting compounds\. Target annotations are dominated by enzymes \(8,163\) and membrane receptors \(5,110\), followed by transporters \(785\), ion channels \(744\), transcription factors, and epigenetic regulators\. The benchmark therefore emphasizes kinases and GPCRs, and generalization to other target classes, such as protein–protein interaction targets, requires further evaluation\.
Figure 8:Benchmark composition\. \(a\) Number of off\-targets per compound on a logarithmic axis, with a median of 3\. \(b\) Clinical stage; approximately two thirds of compounds are approved drugs\. \(c\) Target\-class annotations\. \(d\) Molecular weight versus QED for the starting compounds\.
##### Docking data\.
We attempt to dock each compound against every protein annotated as an intended target or off\-target\. Receptors are resolved per UniProt\([The UniProt Consortium, 2025](https://arxiv.org/html/2609.21165#bib.bib24)\)accession, preferring an experimental PDB structure\([Berman et al\., 2000](https://arxiv.org/html/2609.21165#bib.bib25)\)over an AlphaFold model\([Jumper et al\., 2021](https://arxiv.org/html/2609.21165#bib.bib26)\)and, among PDB candidates, the highest\-resolution holo structure \(one with a bound non\-polymer ligand\) at≤3\.5\\leq 3\.5Å; ligands are embedded with RDKit ETKDGv3\([Wang et al\., 2020](https://arxiv.org/html/2609.21165#bib.bib27)\)and docked with AutoDock Vina\. Both the conformer seed and the Vina seed are fixed to support reproducibility\. We attempted 15,195 compound–protein pairs\. 13,636 returned a score; of those, 143 \(1\.05%\) were physically implausible \(\>−1\>\-1kcal/mol, 117 of them positive, maximum\+43\.9\+43\.9\) and were discarded as failed poses, leaving13,493 usable docking scores: 1,670 against intended targets and 11,823 against off\-targets\. The remaining 1,559 pairs did not return a score, primarily because no protein structure could be resolved\. The full cache is released with the benchmark\.
##### Agreement with measured activity\.
pChEMBL values can represent several endpoints \(KiK\_\{i\},KdK\_\{d\}, IC50, AC50, EC50, and Potency\), so direct correlation with docking scores would combine binding constants with assay\-dependent measures of activity\. To reduce this heterogeneity, we queried ChEMBL at the*activity*level for the 20 most frequent proteins in the benchmark and retained only binding assays \(assay\_type=B\), values reported in nM, and uncensored measurements \(standard\_relation‘=’\)\. For each protein, we selected the endpoint type with the greatest compound coverage, recomputedpp\-activity from the raw values, and summarized replicate measurements by their median\. Coverage ranged from 89–99% of compounds associated with each protein\.
Table[4](https://arxiv.org/html/2609.21165#A1.T4)reports the five proteins with the strongest correlations\. Muscarinic M2 achievesρ=−0\.574\\rho=\-0\.574across 132 compounds, and all five correlations are statistically significant and have the expected negative sign\. These results indicate that docking scores provide information about relative activity for some well\-characterized targets\. However, these five proteins represent the strongest results among 19 evaluated proteins: the overall median isρ=−0\.075\\rho=\-0\.075, and only 6/19 correlations reachp<0\.05p<0\.05\. Agreement is therefore target\-dependent\. Restricting the analysis toKiK\_\{i\}measurements yields negative correlations for 15/18 proteins, of which 7 are statistically significant, with a median correlation of−0\.092\-0\.092\.
Table 4:The five proteins with the strongest docking–affinity agreement, after filtering to a single assay type and unit\. Negativeρ\\rhois the expected direction \(a more negative Vina score should accompany a higherpp\-activity\)\. These are the strongest correlations among the 19 proteins tested; the median over all 19 is−0\.075\-0\.075\.
##### Intended\-target and off\-target score distributions\.
Figure[9](https://arxiv.org/html/2609.21165#A1.F9)groups all usable docking scores according to whether the protein is an intended target or an off\-target of the compound\. Both distributions are unimodal, centered near−8\.2\-8\.2kcal/mol, and span approximately−14\-14to−2\-2kcal/mol\. The two distributions are closely aligned\.
Figure 9:Docking scores pooled into two sets over the whole benchmark: compound–protein pairs where the protein is the compound’s intended target \(n=1,670n=1\{,\}670\) and bindings where it is an off\-target \(n=11,823n=11\{,\}823\)\. Density\-normalized; dashed lines mark medians\. The comparison shows no statistically significant difference and a small effect size: medians−8\.20\-8\.20vs\.−8\.30\-8\.30kcal/mol, Mann–Whitneyp=0\.21p=0\.21, Cohen’sd=0\.02d=0\.02, with quantiles agreeing to within0\.10\.1–0\.40\.4kcal/mol at the 5th, 25th, 75th and 95th percentiles\.
##### Binding\-strength categories\.
We interpret Vina scores as approximate binding free energies to define nominal dissociation\-constant thresholds usingΔG=RTlnKd\\Delta G=RT\\ln K\_\{d\}\(RT=0\.593RT=0\.593kcal/mol at 298 K\)\. Table[5](https://arxiv.org/html/2609.21165#A1.T5)groups all usable docking scores by these thresholds and categorizes experimental measurements using the corresponding pChEMBL cutoffs\.
Table 5:Binding\-strength categories\. Left: our docking scores\. Right: the experimental affinities binned at the pChEMBL equivalents of the same cut points \(−9\.0→6\.59\-9\.0\\to 6\.59,−7\.0→5\.13\-7\.0\\to 5\.13,−5\.0→3\.66\-5\.0\\to 3\.66\)\. Docking places intended targets and off\-targets in nearly identical proportions; the experimental labels differ by approximately 54 percentage points in the strong category\. The experimental “negligible” row is empty because off\-targets enter the benchmark only at pChEMBL≥5\.0\\geq 5\.0\.These results illustrate both the utility and the limitations of the docking scores\. Under the thresholds in Table[5](https://arxiv.org/html/2609.21165#A1.T5), 74\.6% of compound–protein pairs fall within the strong or moderate categories, while 4\.9% fall within the negligible category\. Although this distribution does not independently establish predictive accuracy, the target\-specific correlations in Table[4](https://arxiv.org/html/2609.21165#A1.T4)support the use of docking scores for relative comparisons in some settings\.
However, the scores provide limited separation between intended targets and off\-targets\. Docking assigns similar proportions to the strong or moderate categories \(72\.3% and 74\.7%, respectively\), with differences of at most 2\.6 percentage points in any category\. In contrast, experimental measurements place 81\.5% of intended targets and 27\.7% of off\-targets in the strong category, a difference of approximately 54 percentage points; the corresponding docking difference is only 0\.2 percentage points\. Absolute Vina scores therefore do not reliably distinguish intended targets from off\-targets in this benchmark\. Our optimization evaluation instead uses pre\- and post\-modification*differences*computed with fixed receptors and seeds\. This comparison cancels systematic offsets to the extent that they are shared between the two evaluations, but does not eliminate all scoring errors\.
### A\.3Docking setup
All docking uses AutoDock Vina with receptors and ligands prepared using MGLTools/AutoDockTools\([Morris et al\., 2009](https://arxiv.org/html/2609.21165#bib.bib28)\)\. We report the configuration and implementation safeguards used to address failures observed during large\-scale docking\.
##### Receptor selection\.
ChEMBL target identifiers are mapped to UniProt accessions and then to structures through RCSB PDB\([Berman et al\., 2000](https://arxiv.org/html/2609.21165#bib.bib25)\)\. We query X\-ray or cryo\-EM entries with resolution≤3\.5\\leq 3\.5Å and prefer a holo structure \(one containing at least one bound non\-polymer entity\) over an apo structure with marginally higher resolution\. If no holo candidate is available, we select the highest\-resolution structure\. For example, ranking FLT3 structures solely by resolution selects the apo entry 1RJB \(2\.10 Å\) over the holo entry 6JQR, despite a resolution difference of only0\.10\.1Å\. Prioritizing holo structures accounts for differences in binding\-pocket geometry between ligand\-bound and unbound conformations\.
##### Search box\.
The search box is defined using the following procedures, in order of priority\.
1. 1\.Co\-crystal ligand present\.The box is centered on the bound ligand and sized to its coordinate span plus 12 Å of padding\. Coordinates are taken from the*first occurrence*of the ligand rather than pooled across the structure\. In homo\-oligomeric structures, pooling ligand copies from multiple subunits can expand a pocket\-centered box to encompass the full assembly\. For ALDH2/1O04, this produced a53×161×13153\\times 161\\times 131Å box, exceeding the volume of a typical∼\\sim20 Å pocket box by more than 100\-fold, and increased Vina runtime to 475 s compared with a typical 10–20 s\.
2. 2\.No co\-crystal ligand\.We applyfpocket\([Le Guilloux et al\., 2009](https://arxiv.org/html/2609.21165#bib.bib29)\)to the protein\-only receptor and define the box around the highest\-ranked pocket with 12 Å of padding\. Pockets are ranked by fpocket’s*druggability*score\([Schmidtke and Barril, 2010](https://arxiv.org/html/2609.21165#bib.bib30)\)rather than its default geometric score\. These rankings can differ: for 2ATX, the default first\-ranked pocket had a druggability score of 0\.379, whereas the third\-ranked pocket scored 0\.853\. If no suitable pocket is identified, the fallback is a whole\-protein search box with 10 Å of padding\.
##### Vina parameters\.
\-\-exhaustiveness 8,\-\-cpu 1,\-\-seed 42; the reported score is the top\-ranked mode from the Vina log\. Ligands are embedded with RDKit ETKDGv3\([Wang et al\., 2020](https://arxiv.org/html/2609.21165#bib.bib27)\)and optimized with MMFF\([Halgren, 1996](https://arxiv.org/html/2609.21165#bib.bib31)\)before conversion to PDBQT\.
##### Determinism\.
Our reproducibility configuration fixes three parameters: the Vina seed, the Vina thread count \(\-\-cpu 1\), and the RDKit conformer seed\. The single\-thread setting avoids variation associated with thread scheduling\. RDKit’sEmbedMoleculeuses a defaultrandomSeedof−1\-1, so failing to specify this seed can produce different starting conformers even when the Vina seed is fixed\. Before fixing these parameters, repeated docking of granisetron against the same target produced different binding modes and scores\. With all three parameters fixed, repeated runs produced byte\-identical cache contents\.
##### Concurrency\.
By default, Vina detects all available CPUs \(128 on our system\)\. Concurrent docking calls can therefore oversubscribe computational resources\. With six simultaneous processes, we observed SIGSEGV failures and nonzero exit codes that did not occur when the same commands were executed individually\. We consequently restrict each Vina invocation to one thread and parallelize across independent docking processes\. Each call uses a separate temporary directory to prevent concurrent processes from overwriting intermediate files and corrupting results\.
##### Timeouts\.
Each subprocess call has an explicit timeout: 300 s for Vina and 120 s for MGLTools\. The Python 2 scriptprepare\_ligand4\.pycan fail to terminate for certain ligand topologies; in two observed cases, calls continued at 99\.9% CPU utilization for 70 minutes without producing output\. Becausesubprocess\.runhas no default timeout, explicit limits are necessary to prevent individual ligand\-processing failures from indefinitely occupying workers and progressively reducing available processing capacity\.
##### Contacts\.
For contact\-difference analysis, we extract per\-ligand\-atom contacts using a 7\.0 Å threshold and setmax\_hotspotsto 150\. The underlying routine was designed for fragment growth and prioritizes a small number of high\-volume sites\. We increase the default limit of 5 to obtain broader coverage of the ligand–protein interface\.
##### Caching\.
Scores are cached for each \(compound, protein\) pair in a JSON file shared across the docking campaign\. Writes use a temporary file followed by atomic replacement withos\.replace\. Readers retry after incomplete reads and retain the most recent valid copy\. These safeguards addressJSONDecodeErrorfailures observed when concurrent processes accessed partially written cache files\. The cache is released with the benchmark to enable reproduction of the main analysis without repeating approximately 15,000 docking calls\.
### A\.4Ablation detail
Table[6](https://arxiv.org/html/2609.21165#A1.T6)reports the numerical results corresponding to Figure[6](https://arxiv.org/html/2609.21165#S4.F6)a\. Each row compares the full system with one ablated variant using paired observations from the same compounds\.
Table 6:Paired ablation results\.Δ\\Deltais mean\(full−\-ablated\); positive means the component contributes\.
### A\.5ADMET comparison on the FLT3 setting
Using the same six\-endpoint ADMET oracle, we compare the optimized molecules with both their starting compounds and the reference method’s generated candidates \(Table[7](https://arxiv.org/html/2609.21165#A1.T7)\)\. The optimized molecules have lower predicted AMES, P\-gp, and DILI probabilities and higher predicted HIA than both comparison groups\. In particular, the mean DILI probability is0\.7450\.745, compared with0\.7530\.753for the starting compounds and0\.7870\.787for the generated candidates\. The improvements are not uniform across endpoints: predicted hERG probability increases slightly from0\.6960\.696to0\.7020\.702, while BBBP decreases from0\.6800\.680to0\.6190\.619but remains above the generated candidates’0\.3870\.387\. These results indicate favorable changes in several predicted properties alongside endpoint\-specific trade\-offs\. Because the methods start from different molecular populations and these scores are model predictions rather than experimental measurements, the comparison does not establish overall pharmacological superiority\.
Table 7:ADMET comparison \(mean predicted probability\)\. Arrows indicate the preferred direction:↓\\downarrowlower is better;↑\\uparrowhigher is better\.
### A\.6Implementation details
The complete docking configuration is provided in Section[A\.3](https://arxiv.org/html/2609.21165#A1.SS3)\. Molecule handling uses RDKit\([Landrum, n\.d\.](https://arxiv.org/html/2609.21165#bib.bib11)\), and fragmentation follows BRICS\([Degen et al\., 2008](https://arxiv.org/html/2609.21165#bib.bib10)\)\. The minimum Tanimoto similarity is 0\.4 relative to the original molecule, rather than the molecule from the preceding round, to constrain cumulative structural deviation\. The ADMET tolerance is 0\.15 for AMES, DILI, hERG, and P\-gp\. Unless otherwise specified, runs use 8 proposals per round and 4 rounds\.Similar Articles
ToolMol: Evolutionary Agentic Framework for Multi-objective Drug Discovery
ToolMol is an evolutionary agentic framework that combines a multi-objective genetic algorithm with an LLM-based operator to design small-molecule drugs, achieving state-of-the-art binding affinity and drug-likeness on multiple protein targets.
Probe Before You Edit: Probing-Guided Molecular Optimization for LLM Agents in Structure-Based Drug Design
This paper introduces PROBE, a framework that uses LLM agents to iteratively optimize ligands in structure-based drug design by probing pocket-ligand complex responses before editing, achieving state-of-the-art results on CrossDocked2020.
Generating Developable 3D Molecules via Pocket-Conditioned Diffusion and Property-Aware Optimization
This paper introduces a novel diffusion-based generative model for structure-based drug design that decouples pocket and ligand representation learning and incorporates multi-scale interaction signals and property-aware optimization to generate developable 3D molecules with improved binding affinity and ADMET properties.
Beyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) Benchmark
The Nanotechnology Molecular Optimization (NMO) Benchmark introduces physics-based molecular design tasks replacing drug-discovery-focused metrics, aiming to drive scientific discovery in nanotechnology. The paper shows that advanced methods underperform simpler approaches on NMO, and proposes new baseline methods including a novel representation and domain-agnostic pretraining.
Rethinking Molecular OOD Generalization via Target-Aware Source Selection
This paper introduces SCOPE-Bench, a benchmark for evaluating molecular out-of-distribution generalization, and POMA, a framework using reinforcement learning to select source domains for domain adaptation, achieving significant error reductions on 3D molecular models.