HEDGEHOG: Hierarchical Evaluation of Drug Generators Through Rigorous Filtration

arXiv cs.LG Papers

Summary

Introduces HEDGEHOG, a hierarchical benchmark for evaluating molecular generative models in drug discovery, revealing that only 0.65% of generated molecules pass all medicinal chemistry and docking filters.

arXiv:2607.13155v1 Announce Type: new Abstract: Generative molecular models can support early drug discovery by proposing new candidate compounds de novo. In practice, useful candidates must balance target-relevant activity, synthetic accessibility, physicochemical properties, and other multiparameter design constraints. However, metrics commonly used to evaluate molecular generators only weakly reflect whether the generated compounds are medicinally plausible and suitable for downstream computation. This can produce false positives in model evaluation, incorrect assumptions, and inefficient use of computational resources. We introduce HEDGEHOG, a unified six-stage filtration benchmark that is inspired by industrial hit identification workflows: (i) preprocessing; (ii) physicochemical descriptor screening; (iii) structural alerts and graph-sanity checks; (iv) synthesis feasibility; (v) docking and binding affinity estimation; and (vi) three-dimensional pose and interaction checks. We evaluate 23 molecular generators across three model classes under a standardized protocol. Across 230,000 generated molecules, only 0.65% of initial molecules survive all stages. Our results expose a central limitation of current molecular generators: molecules that appear acceptable under isolated criteria rarely satisfy medicinal chemistry, synthesis, docking, and 3D pose filters simultaneously.
Original Article
View Cached Full Text

Cached at: 07/16/26, 04:21 AM

# HEDGEHOG: Hierarchical Evaluation of Drug Generators Through Rigorous Filtration
Source: [https://arxiv.org/html/2607.13155](https://arxiv.org/html/2607.13155)
###### Abstract

Generative molecular models can support early drug discovery by proposing new candidate compoundsde novo\. In practice, useful candidates must balance target\-relevant activity, synthetic accessibility, physicochemical properties, and other multiparameter design constraints\. However, metrics commonly used to evaluate molecular generators only weakly reflect whether the generated compounds are medicinally plausible and suitable for downstream computation\. This can produce false positives in model evaluation, incorrect assumptions, and inefficient use of computational resources\. We introduceHEDGEHOG, a unifiedsix\-stagefiltration benchmark that is inspired by industrial hit identification workflows: \(i\) preprocessing; \(ii\) physicochemical descriptor screening; \(iii\) structural alerts and graph\-sanity checks; \(iv\) synthesis feasibility; \(v\) docking and binding affinity estimation; and \(vi\) three\-dimensional pose and interaction checks\. We evaluate2323molecular generators across three model classes under a standardized protocol\. Across230,000230,000generated molecules, only0\.65%0\.65\\%of initial molecules survive all stages\. Our results expose a central limitation of current molecular generators: molecules that appear acceptable under isolated criteria rarely satisfy medicinal chemistry, synthesis, docking, and 3D pose filters simultaneously\.

## 1Introduction

Generative models are widely used in molecular design for drug discovery, enabling the generation of large numbers of candidate compounds through efficient exploration and exploitation of chemical space\[Tropsha et al\.,[2024](https://arxiv.org/html/2607.13155#bib.bib74)\]\. However, the ability to generate chemically valid molecules alone is insufficient for practical drug development\[Ivanenkov et al\.,[2023](https://arxiv.org/html/2607.13155#bib.bib33); Du et al\.,[2024](https://arxiv.org/html/2607.13155#bib.bib15)\]\. To serve as meaningful starting points for medicinal chemistry, generated molecules must satisfy multiple criteria related to drug\-likeness, developability, and biological relevance\[Atz et al\.,[2024](https://arxiv.org/html/2607.13155#bib.bib2)\]\.

Standard evaluations of molecular generators emphasize intrinsic metrics such as validity, novelty, uniqueness, and distributional similarity\[Brown et al\.,[2019](https://arxiv.org/html/2607.13155#bib.bib11)\]\. These metrics are useful for assessing whether a model produces chemically plausible and nontrivial outputs; however, they are not sufficient on their own to establish practical value in early\-stage drug discovery\. For those applications, compounds are prioritized using multiple medicinal chemistry and drug discovery criteria beyond potency alone, including physicochemical properties\[Veber et al\.,[2002](https://arxiv.org/html/2607.13155#bib.bib77)\], structural liabilities\[Bruns and Watson,[2012](https://arxiv.org/html/2607.13155#bib.bib12)\], synthetic tractability\[Du et al\.,[2024](https://arxiv.org/html/2607.13155#bib.bib15); Swanson et al\.,[2024](https://arxiv.org/html/2607.13155#bib.bib71); Segler et al\.,[2018](https://arxiv.org/html/2607.13155#bib.bib68)\], and structure\-based screening procedures\[Zhou et al\.,[2024](https://arxiv.org/html/2607.13155#bib.bib86); Zhao,[2024](https://arxiv.org/html/2607.13155#bib.bib84)\]\.

This gap between intrinsic generation quality and practical utility is especially important in hit identification\. Early discovery campaigns do not operate on raw strings or abstract graph distributions\. Instead, they use molecules that balance multiple properties\[Duffy et al\.,[2012](https://arxiv.org/html/2607.13155#bib.bib16)\]\. Later analysis stages are more time\- and resource\-intensive and often constrained by limited downstream compute budgets\. Therefore, these stages are applied selectively to a prescreened set of molecules rather than exhaustively on the full generated set\[Dahlin and Walters,[2014](https://arxiv.org/html/2607.13155#bib.bib13); Bedart et al\.,[2024](https://arxiv.org/html/2607.13155#bib.bib6)\]\. Molecules with implausible properties, reactive or unstable motifs, poor synthetic tractability, or incompatible binding poses consume resources without improving the quality of the process\. As a result, evaluating generators only with broad distributional or structural metrics can substantially overestimate their usefulness for medicinal chemistry\.

Several task\-specific benchmarks have incorporated docking scores, property objectives, or synthetic accessibility proxies\. However, these objectives are often evaluated individually or in simplified combinations, and they rarely reproduce the sequential attrition imposed by practical hit identification workflows\. As a result, a model may appear successful by optimizing a property, pharmacophore, or docking proxy while still producing molecules that fail basic medicinal chemistry, retrosynthesis, or pose plausibility checks\.

To address these limitations, we developHEDGEHOG, a multi\-stage benchmark designed to evaluate generated molecules under a computational triage workflow that more closely reflects hit identification practice\. The benchmark applies a fixed sequence of filters covering preprocessing, descriptor screening, structural alerts, synthesis feasibility, docking and affinity estimation, and post\-docking three\-dimensional \(3D\) checks\. Rather than relying on common generative metrics,HEDGEHOGmeasures whether model outputs remain plausible after running medicinal\-chemistry\-oriented criteria\.

This work provides a unified protocol that can be applied to unconditional, ligand\-based, and protein\-based models, enabling stage\-wise comparison under the same downstream constraints\. We instantiateHEDGEHOGon molecules targeting the switch\-II pocket of KRAS G12D, a therapeutically important and structurally challenging target\[Mao et al\.,[2022](https://arxiv.org/html/2607.13155#bib.bib50); Vasta et al\.,[2022](https://arxiv.org/html/2607.13155#bib.bib76)\]\. In addition to cumulative survival, we report conditional survival at each stage in order to identify where a model class loses viability as the screening procedure becomes more stringent\. This framing makes it possible to separate strong performance on conventional intrinsic metrics from performance under a more demanding computational triage setting\.

These results establishHEDGEHOGas a stress test for molecular generators, exposing failure modes that are not captured by conventional intrinsic metrics\. The proposed benchmark reframes molecular generation evaluation as a problem of useful chemical space proposal rather than unconstrained molecule generation\. By identifying where generated molecules fail during a realistic triage workflow,HEDGEHOGprovides a more informative basis for developing molecular generative models that are better aligned with early drug discovery\.

## 2Related Work

Distribution\-learning benchmarks\.Existing benchmark suites for molecular generation largely assess two\-dimensional \(2D\) molecular quality and how closely generated sets reproduce reference distributions\. GuacaMol\[Brown et al\.,[2019](https://arxiv.org/html/2607.13155#bib.bib11)\]provides standardized tasks for both distribution learning and goal\-directed optimization, reporting metrics such as validity, uniqueness, novelty, KL divergence, and Fréchet ChemNet Distance\[Preuer et al\.,[2018](https://arxiv.org/html/2607.13155#bib.bib64)\]\. It incorporates compound quality metrics using rule\-based filters derived from SureChEMBL, Glaxo\[Hann et al\.,[1999](https://arxiv.org/html/2607.13155#bib.bib28)\], PAINS\[Baell and Holloway,[2010](https://arxiv.org/html/2607.13155#bib.bib3)\], and in\-house rule sets\. MOSES\[Polykovskiy et al\.,[2020](https://arxiv.org/html/2607.13155#bib.bib63)\]standardizes dataset splits and intrinsic distributional metrics\. It reports the fraction of molecules passing medicinal chemistry filters such as MCF and PAINS, and evaluates SA score, although this score is a heuristic proxy for synthesis feasibility\. These benchmarks do not include comprehensive synthetic feasibility estimation, docking and binding pose assessment, or 3D structure\-dependent tests\.

Structure\-based benchmarks\.Recent benchmarks move closer to practical structure\-based evaluation\. DOCKSTRING\[García\-Ortegón et al\.,[2022](https://arxiv.org/html/2607.13155#bib.bib21)\]packages docking into a reproducible benchmark with a large target panel and tasks such as virtual screening andde novodocking score optimization\. DOCKSTRING also reports a small set of molecular descriptors, including HBDs, HBAs, rotatable bond count, and logP\. GenBench3D\[Baillif et al\.,[2024](https://arxiv.org/html/2607.13155#bib.bib5)\]evaluates 3D molecular generators inside binding pockets, introduces geometry\-aware Validity3D criteria for conformational plausibility, and reports docking\-based scores from Vina\[Trott and Olson,[2010](https://arxiv.org/html/2607.13155#bib.bib75)\], Glide\[Friesner et al\.,[2004](https://arxiv.org/html/2607.13155#bib.bib19)\], and Gold PLP\[Verdonk et al\.,[2003](https://arxiv.org/html/2607.13155#bib.bib78)\]\. It computes a narrow panel of molecular descriptors, including molecular weight, logP, and ring proportion, and uses SA score as a proxy for synthesizability\. Durian\[Nie et al\.,[2024](https://arxiv.org/html/2607.13155#bib.bib55)\]further expands structure\-based 3D evaluation by incorporating protein–ligand complexes with experimental affinity annotations\. It uses a broader panel of physicochemical and geometric metrics, and evaluates docking\-based affinity with QuickVina2\[Alhossary et al\.,[2015](https://arxiv.org/html/2607.13155#bib.bib1)\], Surflex\[Jain,[2003](https://arxiv.org/html/2607.13155#bib.bib35)\], and GNINA\[McNutt et al\.,[2021](https://arxiv.org/html/2607.13155#bib.bib53)\]\. However, these benchmarks do not natively integrate synthesizability estimation, structural alert filtering, or a representative descriptor panel\.

Synthesis\-aware benchmarks\.Synthesis\-aware evaluation has also received increasing attention\. SDDBench\[Liu et al\.,[2024](https://arxiv.org/html/2607.13155#bib.bib47)\]argues that standard synthetic\-accessibility scores are insufficient and instead evaluates molecules using a retrosynthesis\-based criterion together with search success rate\. TARTARUS\[Nigam et al\.,[2023](https://arxiv.org/html/2607.13155#bib.bib56)\]introduces practical objectives based on physical simulation and explicit computational budgets\. Its drug design tasks combine docking objectives with structural constraints and medicinal chemistry filters, while also incorporating metrics such as QED, TPSA, and SA score\. These works move evaluation closer to downstream utility, but they are not designed for efficient reduction to actionable chemical space\.

General\-purpose evaluation framework\.Frameworks such as MolScore\[Thomas et al\.,[2024](https://arxiv.org/html/2607.13155#bib.bib73)\]and TDC\[Huang et al\.,[2021](https://arxiv.org/html/2607.13155#bib.bib31)\]provide flexible infrastructure for composing objectives, datasets, and evaluation modules\. They are valuable for building and testing molecular design workflows, and in some cases contain many of the components that a practical pipeline would use\. Our aim is complementary\. Rather than proposing another general\-purpose scoring framework, we define a highly configurable benchmarking protocol with a fixed staged structure and stage\-wise reporting, motivated by hit identification triage\.HEDGEHOGevaluates whether generated molecules continue to survive under a sequence of increasingly restrictive and practically motivated filters\. Table[1](https://arxiv.org/html/2607.13155#S2.T1)summarizes the relationship between prior work andHEDGEHOG\.

Table 1:Comparison ofHEDGEHOGwith prior benchmarks and evaluation frameworks for molecular generation\. Each column shows whether a benchmark covers one stage of theHEDGEHOGworkflow\. “Yes” means that the stage is evaluated directly or supported out of the box, “Partial” means that only part of the stage is covered through a proxy metric or narrower components, and “No” means that the stage is not covered\. “Code” indicates whether an official implementation or public repository is availableHEDGEHOG benchmark protocol\.HEDGEHOGevaluates model outputs by whether compounds pass a staged filtration cascade that reflects how medicinal chemistry teams triage molecules before hit identification\. It analyzes how different generator classes fail under fixed downstream constraints\. Beyond standard intrinsic metrics,HEDGEHOGcomputes2121physicochemical descriptors, and applies structural alerts and graph sanity checks using approximately2,4602,460SMARTS patterns and1010rule sets\. It assesses synthesizability using four criteria together with explicit synthesis route finding in AiZynthFinderGenheden et al\. \[[2020](https://arxiv.org/html/2607.13155#bib.bib23)\], and evaluates docking usingsmina\[Koes et al\.,[2013](https://arxiv.org/html/2607.13155#bib.bib41)\], GNINA\[McNutt et al\.,[2021](https://arxiv.org/html/2607.13155#bib.bib53)\], and Matcha\[Frolova et al\.,[2025](https://arxiv.org/html/2607.13155#bib.bib20)\]alongside binding affinity prediction using Boltz\-2\[Passaro et al\.,[2025](https://arxiv.org/html/2607.13155#bib.bib60)\]\. Finally, our benchmark applies 3D molecular filters, including pose quality, conformer deviation, similarity, and interaction\-based assessments\. Appendix Table[4](https://arxiv.org/html/2607.13155#A1.T4)contains software versions for each tool\.

## 3Methods

HEDGEHOG\(Fig\.[1](https://arxiv.org/html/2607.13155#S4.F1)\) comprises six stages: molecular preprocessing, physicochemical filtering, structural filtering, synthesis feasibility assessment, docking and binding affinity estimation, and 3D post\-docking analysis\. The staged design is motivated not only by medicinal chemistry practice but also by computational efficiency: retrosynthetic route finding, docking, and binding affinity estimation are substantially more expensive than early descriptor\- and structure\-based checks\.HEDGEHOGtherefore uses a coarse\-to\-fine filtration cascade that removes implausible molecules before sending a much smaller candidate set to the slowest downstream stages\. The overall thresholds and rule sets are target\-agnostic\. In this study, the benchmark is instantiated for KRAS G12D at the switch\-II pocket \(PDB ID: pdb\_00007ew9\[Shen et al\.,[2021](https://arxiv.org/html/2607.13155#bib.bib69)\]\)\. Full configuration details per stage and thresholds are provided in Appendix Section[A](https://arxiv.org/html/2607.13155#A1)\.

### 3\.1Stage 1: Molecular preprocessing

All generated molecules are first passed through a common preprocessing procedure implemented with RDKit\[Landrum,[2013](https://arxiv.org/html/2607.13155#bib.bib44)\]and DataMol\[Mary et al\.,[2024](https://arxiv.org/html/2607.13155#bib.bib51)\]\. This stage converts generated strings into molecular objects, removes salts and solvents, retains the largest relevant organic fragment, disconnects metals, standardizes representation through normalization and sanitization, and converts back to SMILES strings\. Molecules containing atom types outside the supported element set, radicals, isotopes, invalid valence states, or unresolved multi\-fragment artifacts are removed at this stage\. This preprocessing serves two purposes\. First, it standardizes outputs from heterogeneous generators before evaluation\. Second, it reduces the risk that downstream calculations are affected by malformed or chemically inconsistent inputs \(Appendix Section[A\.1](https://arxiv.org/html/2607.13155#A1.SS1)\)\.

### 3\.2Stage 2: Physicochemical descriptors

After preprocessing,HEDGEHOGcomputes a panel of 21 physicochemical descriptors and filters molecules using configurable threshold ranges \(Appendix Table[5](https://arxiv.org/html/2607.13155#A1.T5)\)\. The panel includes size\- and composition\-related descriptors, lipophilicity and polarity descriptors, hydrogen bond features, ring statistics, fraction of sp3\-hybridized carbon atoms, and additional medicinal chemistry checks\. We do not adopt a single canonical rule set like Lipinski\[Lipinski,[2004](https://arxiv.org/html/2607.13155#bib.bib45)\]or Veber\[Veber et al\.,[2002](https://arxiv.org/html/2607.13155#bib.bib77)\], because none of them covers the full descriptor panel used in our benchmark, and published thresholds vary across target classes, screening settings, and medicinal chemistry objectives\. Instead, we use broad descriptor bounds that combine literature\-reported thresholds with manually curated checks for descriptors where no widely accepted cutoff exists\. These thresholds are chosen to remove extreme outliers and clearly implausible molecules while preserving broad chemical diversity\.

### 3\.3Stage 3: Structural filters

Stage 3 removes molecules with undesirable substructures or poor medicinal chemistry profiles\. This stage combines public structural alert libraries with additional rule\-based graph checks to remove chemotypes that are excluded in small molecule hit identification workflows, including unstable, highly reactive, toxic, and assay\-interfering motifs, as well as structures with inappropriate graph\-level features\. The alert layer covers medicinal chemistry and assay\-interference substructures such as Dundee\[Brenk et al\.,[2008](https://arxiv.org/html/2607.13155#bib.bib10)\], BMS\[Pearce et al\.,[2006](https://arxiv.org/html/2607.13155#bib.bib61)\], Glaxo\[Hann et al\.,[1999](https://arxiv.org/html/2607.13155#bib.bib28)\], PAINS\[Baell and Holloway,[2010](https://arxiv.org/html/2607.13155#bib.bib3)\], Lilly Medchem Rules\[Bruns and Watson,[2012](https://arxiv.org/html/2607.13155#bib.bib12)\], NIBR\[Schuffenhauer et al\.,[2020](https://arxiv.org/html/2607.13155#bib.bib67)\], and related alert sets \(Appendix Table[6](https://arxiv.org/html/2607.13155#A1.T6)\)\. The rule\-based layer includes graph\-level checks for ring infractions, protecting groups, halogen\- and stereochemistry\-related constraints, and other structural features that are typically excluded during medicinal chemistry review\. Because some alert categories can be context\-dependent in broader drug discovery settings, the structural filters stage is fully configurable, both at the level of rule set selection and at the level of individual SMARTS patterns \(Supplementary File 2\)\.

### 3\.4Stage 4: Synthesis feasibility

We evaluate whether molecules are likely to be synthetically tractable by combining heuristic synthesis\-related scores with explicit retrosynthesis search\. In the default configuration,HEDGEHOGcomputes SA score\[Ertl and Schuffenhauer,[2009](https://arxiv.org/html/2607.13155#bib.bib17)\], RA score\[Thakkar et al\.,[2021](https://arxiv.org/html/2607.13155#bib.bib72)\], and SYBA score\[Voršilák et al\.,[2020](https://arxiv.org/html/2607.13155#bib.bib79)\], and then runs AiZynthFinder\[Genheden et al\.,[2020](https://arxiv.org/html/2607.13155#bib.bib23)\]to test whether at least one synthetic route can be found under the chosen retrosynthesis setup\. Because retrosynthetic search is computationally expensive, it is applied only after faster physicochemical and structural filters have substantially reduced the candidate pool\. The default thresholds for synthesis scores are SA≤4\.5\\leq 4\.5, SYBA≥0\\geq 0, and RA≥0\.5\\geq 0\.5\. We use this combination because heuristic scores alone are insufficient\. They are fast and informative; however, they do not provide an explicit route\. At the same time, route finding success depends on the template library, stock set, search policy, and compute budget\. An example of a synthetic route provided by AiZynthFinder is shown in Appendix Fig\.[5](https://arxiv.org/html/2607.13155#A1.F5)\.

### 3\.5Stage 5: Docking and binding affinity estimation

Molecules that pass the earlier stages are evaluated in a target\-aware manner by docking them into the KRAS G12D switch\-II pocket\.HEDGEHOGsupports several docking engines, includingsmina\[Koes et al\.,[2013](https://arxiv.org/html/2607.13155#bib.bib41)\], GNINA\[McNutt et al\.,[2021](https://arxiv.org/html/2607.13155#bib.bib53)\], and Matcha\[Frolova et al\.,[2025](https://arxiv.org/html/2607.13155#bib.bib20)\], a recent neural docking method based on multi\-stage Riemannian flow matching\. These tools were selected based on their strong performance in independent evaluations on diverse docking datasets\[Pak et al\.,[2025](https://arxiv.org/html/2607.13155#bib.bib59)\]\. In the default configuration,HEDGEHOGuses all three docking tools together and estimates binding affinity with Boltz\-2\[Passaro et al\.,[2025](https://arxiv.org/html/2607.13155#bib.bib60)\]\. Docking and affinity estimation are among the slowest stages in the pipeline, so they are performed only on molecules that survive the earlier low\-cost filters\. Because docking scores and affinity estimates are imperfect proxies for true binding, we extend this stage with multiple models, rather than relying on any single proxy\. Survival through Stage 5 is a convergent computational support for favorable predicted activity and plausible binding\.

Molecules that reach docking have already passed validity, descriptor, structural, and synthesis feasibility checks\. These early stages are essential because docking scores alone can be exploited by unrealistic molecules \(Appendix Fig\.[6](https://arxiv.org/html/2607.13155#A1.F6)\)\. A detailed explanation and the corresponding thresholds are provided in Appendix Section[A\.5](https://arxiv.org/html/2607.13155#A1.SS5)\.

### 3\.6Stage 6: Three\-dimensional filtration

The final stage evaluates docked poses using a sequence of post\-docking 3D checks\. These checks are intended to remove molecules whose nominal docking success is undermined by poor poses or implausible geometry\. That is why we include search\-box containment, pose quality assessment, protein\-ligand interaction checks, and a conformer\-deviation analysis\. For the KRAS G12D benchmark, we impose a target\-specific interaction requirement involving Asp12\. The interaction is illustrated in Appendix Fig\.[7](https://arxiv.org/html/2607.13155#A1.F7)\.HEDGEHOGseparates target\-agnostic 3D checks from target\-specific interaction rules so that future benchmark instances can define different target\-aware criteria without changing the broader framework\. The interaction stage can enforce required or forbidden residue contacts through ProLIF\-based analysis\[Bouysset and Fiorucci,[2021](https://arxiv.org/html/2607.13155#bib.bib9)\]\.

## 4Results and discussion

We evaluate three classes of molecular generators, includingunconditional,ligand\-based, andprotein\-basedmodels\. For each model, we sampleNg​e​n=10,000N\_\{gen\}=10,000distinct SMILES strings\. Invalid strings were not included in the10,00010,000SMILES input set\. During this step, uniqueness was enforced only by exact string matching\. Therefore, different SMILES strings encoding the same molecular graph could still be retained\. However, during Stage 1 \(molecular preprocessing\), redundant representations of generated molecules may be reduced through SMILES sanitization and standardization of structures, applying structure\-level deduplication, and removal of chemically inconsistent molecules\. For modelmmand stagess, survival with respect to initial number of molecules was defined asNm,s/Nm,0N\_\{m,s\}/N\_\{m,0\}, whereNm,0=10,000N\_\{m,0\}=10\{,\}000\. Survival with respect to the previous stage was defined asNm,s/Nm,s−1N\_\{m,s\}/N\_\{m,s\-1\}\. For model level summaries, we report the mean and standard deviation of the number of retained molecules across individual generators within each model class\. All experiments were conducted using an NVIDIA A100 40 GB GPU\. Runtime details are provided in Appendix Section[A\.7](https://arxiv.org/html/2607.13155#A1.SS7)\.

### 4\.1HEDGEHOGbenchmark design and evaluated generators

![Refer to caption](https://arxiv.org/html/2607.13155v1/img/Hedgehog.png)Figure 1:HEDGEHOGevaluates generated molecules with a six\-stage, coarse\-to\-fine filtering workflow\. Starting from input SMILES, the pipeline first cleans and standardizes each molecule, then applies physicochemical filters, structural and medicinal chemistry filters, and synthesis feasibility filters before running the more expensive docking and binding affinity, and final 3D post\-docking checks\.HEDGEHOGis a hierarchical filtration benchmark designed to evaluate molecular generators under a staged computational workflow\. Generated molecules are first evaluated with inexpensive, broadly applicable medicinal chemistry filters and only then passed to more expensive target\-aware calculations\. The complete workflow is shown in Fig\.[1](https://arxiv.org/html/2607.13155#S4.F1), and full implementation details and thresholds are provided in Methods and Appendix Section[A](https://arxiv.org/html/2607.13155#A1)\.

HEDGEHOGapplies six sequential filters to the generated molecules\. First, SMILES strings are standardized and chemically invalid structures, unsupported atoms, salts, solvents, and disconnected fragments are removed\. Second, molecules are screened by 21 physicochemical descriptors that capture composition, lipophilicity, polarity, and molecular complexity\[Ivanenkov et al\.,[2019](https://arxiv.org/html/2607.13155#bib.bib34)\]\. Third, undesirable structural motifs are removed using public alert sets such as PAINS\[Baell and Holloway,[2010](https://arxiv.org/html/2607.13155#bib.bib3)\], Lilly Medchem Rules\[Bruns and Watson,[2012](https://arxiv.org/html/2607.13155#bib.bib12)\]and NIBR\[Schuffenhauer et al\.,[2020](https://arxiv.org/html/2607.13155#bib.bib67)\]\. Fourth, synthesis feasibility is assessed with SA\[Ertl and Schuffenhauer,[2009](https://arxiv.org/html/2607.13155#bib.bib17)\], RA\[Thakkar et al\.,[2021](https://arxiv.org/html/2607.13155#bib.bib72)\]and SYBA\[Voršilák et al\.,[2020](https://arxiv.org/html/2607.13155#bib.bib79)\]scores and explicit route search in AiZynthFinder\[Genheden et al\.,[2020](https://arxiv.org/html/2607.13155#bib.bib23)\]\. Fifth, retained molecules are docked into the target protein withsmina\[Koes et al\.,[2013](https://arxiv.org/html/2607.13155#bib.bib41)\], GNINA\[McNutt et al\.,[2021](https://arxiv.org/html/2607.13155#bib.bib53)\]and Matcha\[Frolova et al\.,[2025](https://arxiv.org/html/2607.13155#bib.bib20)\], and evaluated by Boltz\-2\[Passaro et al\.,[2025](https://arxiv.org/html/2607.13155#bib.bib60)\]affinity prediction\. Finally, docked poses are checked for 3D plausibility, including pocket containment, pose quality, conformer deviation and target\-specific interaction\.

We evaluated2323molecular generators grouped into three classes ofunconditional,ligand\-based, andprotein\-basedmodels\. For each model, we generated10,00010,000valid and distinct SMILES, excluding unparsable strings, and with uniqueness enforced at the raw string level\. This yielded80,00080,000molecules from unconditional models,70,00070,000molecules from ligand\-based models, and80,00080,000molecules from protein\-based models, for a total of230,000230,000generated molecules\. Evaluation details are provided in Methods section[3](https://arxiv.org/html/2607.13155#S3)\.

### 4\.2Survival under theHEDGEHOGworkflow

We first quantified how many generated molecules survive the complete filtration cascade and at which stages attrition occurs\. Table[2](https://arxiv.org/html/2607.13155#S4.T2)summarizes outputs at each stage for each model class, and reports both survival relative to the initial generated set and survival relative to the previous stage\. This distinction is important because the lowest final survival is not always caused by the final target\-aware filters\. The per model statistics provide a complementary view of these class\-level trends by reporting the mean and standard deviation of the number of retained molecules across individual generators within each model class\. Since each generator contributes10,00010,000valid and distinct SMILES strings, these values quantify both the typical survival of a single generator and the dispersion of performance within the class\. Fig\.[2](https://arxiv.org/html/2607.13155#S4.F2)reports cumulative survivors after each stage\.

Table 2:Stage\-wise attrition of generated molecules by model class\. Counts report the cumulative number of molecules retained after eachHEDGEHOGstage\. Unconditional and protein\-based models each start from 80,000 molecules, and ligand\-based models start from 70,000 molecules\. For each model class, the first percentage column reports cumulative survival relative to the initial set, and the second percentage column reports survival relative to previous stage\. The per model column reports the number of molecules as the mean and standard deviation, number of retained molecules across individual generators within the corresponding model classEarly\-stage filtering\.Preprocessing already shows differences in raw generation quality\. Conditional models produce molecules that pass standardization at much higher rates than unconditional models, with average preprocessing pass rates97\.56%97\.56\\%and75\.51%75\.51\\%, respectively\. At preprocessing, unconditional models retain7551±34767551\\pm 3476molecules per generator, whereas ligand\-based and protein\-based models retain9837±1969837\\pm 196and9675±7439675\\pm 743, respectively\. Thus, conditional models improve early standardization not only in aggregate, but also with greater consistency across generators\. As shown in Fig\.[2](https://arxiv.org/html/2607.13155#S4.F2), several unconditional generators, such as E\(3\)DM and TGM\-DLM, lose a substantial fraction of molecules before any chemical filtering is applied, highlighting poor raw generation quality\.

\\phantomsubcaption

a

![Refer to caption](https://arxiv.org/html/2607.13155v1/img/survival_plots/survival_plot_un.png)

\\phantomsubcaption

b

![Refer to caption](https://arxiv.org/html/2607.13155v1/img/survival_plots/survival_plot_lb.png)

\\phantomsubcaption

c

![Refer to caption](https://arxiv.org/html/2607.13155v1/img/survival_plots/survival_plot_pb.png)

Figure 2:Survival of generated molecules through sequential molecular\-design filters\.a–c, Stepwise survival curves showing the percentage of 10,000 generated SMILES strings retained after each filtering stage for \(a\) unconditional, \(b\) ligand\-based, and \(c\) protein\-based molecular generation models\. Each curve tracks the remaining fraction of molecules after preprocessing, descriptor\-based filtering, structural filters, synthetic\-feasibility assessment, docking\- and binding\-affinity filtering, and final three\-dimensional filtering\. Legends report the final number of molecules retained by each model\. Endpoint annotations highlight the highest final survival rates\.Descriptor filtering is the first major bottleneck in the benchmark\. High retention rates on individual filters do not guarantee simultaneous satisfaction, which is substantially more difficult\. Therefore, the main challenge is multi\-property balance rather than isolated property failure\. According to Table[2](https://arxiv.org/html/2607.13155#S4.T2), unconditional, ligand\-based, and protein\-based model classes are reduced to smaller sets of24\.93%24\.93\\%,28\.54%28\.54\\%, and24\.27%24\.27\\%respectively\. This convergence at the class level masks substantial variation between individual generators\. The corresponding per model values are2493±20622493\\pm 2062,2854±21092854\\pm 2109, and2427±14232427\\pm 1423molecules for unconditional, ligand\-based, and protein\-based models, respectively\. The large standard deviations indicate that descriptor compliance is highly model\-dependent, even within the same generation setting\. Ligand\-based models perform slightly better, but the variation is large, for example, GCPG and GENTRL retain 5,849 and 5,451 molecules respectively, while PGMG drops to 97 molecules\. Structural filters remove an additional fraction of molecules\. Unconditional, ligand\-based, and protein\-based model classes fall to5\.82%5\.82\\%,5\.90%5\.90\\%, and3\.62%3\.62\\%of their initial molecules, respectively\. This suggests that many molecules that satisfy a broad panel of physicochemical descriptors still contain prohibited substructures\. Without a structural evaluation stage, models may appear competitive in later non\-structural filtering stages while systematically enriching undesirable chemotypes\. Since the descriptor and structural filtering stages are target\-agnostic and rely on simple, broadly applicable criteria, the fact that only5\.11%5\.11\\%of all initial molecules remain for synthesis evaluation suggests that many generated compounds fail general standards of chemical plausibility and developability\. This implies that the main weakness of most models is not only in later proxy\-based optimization, but already in their ability to generate molecules that satisfy common baseline requirements for drug\-like chemical space\. The staged design ofHEDGEHOGalso serves a practical purpose as a coarse\-to\-fine computational cascade\. By removing most candidates before expensive downstream analyses, these early stages reduce the number of molecules entering synthesis feasibility and docking evaluation, thereby lowering overall computational time and cost\.

Synthesis evaluation and docking\-based filtering\.Explicit retrosynthetic planning is substantially more selective than heuristic synthesizability scores alone\. After synthesis feasibility filtering, unconditional, ligand\-based, and protein\-based model classes retained3\.47%3\.47\\%,2\.12%2\.12\\%, and1\.65%1\.65\\%of their initial molecules, respectively\. This shows that satisfying score\-based proxies alone can overestimate practical synthetic accessibility\. Conditional pass rates show that synthesis filtering is not equally severe for all classes\. Among molecules that passed structural filtering, unconditional, protein\-based, and ligand\-based models retained59\.72%59\.72\\%,45\.44%45\.44\\%, and35\.89%35\.89\\%, respectively\. Thus, ligand\-based or protein\-based generation did not necessarily improve robustness under broader synthetic tractability constraints\. The remaining molecules were then evaluated by docking and binding affinity filters\. Relative to the initial set, unconditional models retained the highest fraction of molecules, compared with ligand\-based and protein\-based models, retaining1\.80%1\.80\\%,1\.55%1\.55\\%, and0\.96%0\.96\\%, respectively\. As shown in Supplementary File 1, docking with a single tool appears to result in tool\-specific false positive pass rates\. The bottleneck at this stage is not strong performance in any single docking or affinity model, but agreement across models with different inductive biases\. Molecules that satisfy score\-based docking criteria may still fail geometric or conformational plausibility checks\. The final 3D filters assess whether each retained docked pose is geometrically plausible, conformationally accessible, and consistent with the target\-specific interaction hypothesis\.

Model\-class performance overview\.Unconditional models retain the highest fraction of molecules through the docking and binding\-affinity stage relative to their initial set \(1\.80%1\.80\\%\), compared with ligand\-based models \(1\.55%1\.55\\%\) and protein\-based models \(0\.96%0\.96\\%\)\. This result should be interpreted after the preceding target\-agnostic filters\. UnderHEDGEHOG, where docking is evaluated only after these chemically implausible molecules are removed, explicit target conditioning does not lead to higher survival at the docking and binding affinity estimation stage\.

Ligand\-based models show the highest conditional survival from the synthesis feasibility stage to the docking and binding affinity stage, retaining73\.10%73\.10\\%of synthesis stage survival molecules\. However, unconditional models retain the largest absolute number of molecules at this stage because they enter it with a larger post\-synthesis pool\. Therefore, the lower end\-to\-end survival of ligand\-based models \(Fig\.[2](https://arxiv.org/html/2607.13155#S4.F2)\) is the combined effect of earlier losses\.

Among model classes, protein\-based generators have the lowest cumulative survival through docking and affinity estimation \(Fig\.[2](https://arxiv.org/html/2607.13155#S4.F2)\), yet the highest conditional survival at the 3D stage, retaining63\.15%63\.15\\%of docking survivors\. This suggests that, once protein\-conditioned molecules pass the earlier filters and docking affinity criteria, their retained poses are more likely to satisfy the final geometric and interaction\-based constraints\. However, because fewer protein\-based molecules survive the earlier target\-agnostic stages, this local advantage does not translate into the highest end\-to\-end survival\.

The last stage reduces the benchmark to a small set of molecules that satisfy allHEDGEHOGfilters\. Unconditional, ligand\-based, and protein\-based model classes retain609609,396396, and485485molecules in total, corresponding to0\.76%0\.76\\%,0\.57%0\.57\\%, and0\.61%0\.61\\%, respectively, suggesting that target conditioning alone does not guarantee higher end\-to\-end survival\. Protein\-based models retain61±11061\\pm 110, where the standard deviation is substantially larger than the mean\. This means that the protein\-based class total is driven disproportionately by the strongest Dragonfly model rather than by uniformly strong performance across the class\.

HEDGEHOGselects for a narrow combination of valid generation, medicinal chemistry quality, synthetic accessibility, binding tool agreement, and physically plausible 3D poses\. This explains why models with high early\-stage retention do not necessarily produce the best final candidates\. The differences between ligand\-based REINVENT4 configurations show that the broad out\-of\-the\-box prior outperforms both the similarity\-biased and transfer\-learned versions, suggesting that stronger target bias can overconstrain generation and reduce robustness under broad medicinal chemistry filters\.

Table 3:FinalHEDGEHOGsurvivors by generator within each model class, ranked by the number of molecules that pass the full pipelineTo provide a compact cross\-class summary, Table[3](https://arxiv.org/html/2607.13155#S4.T3)reports the best performing generators within each model class, ranked by the number of molecules that pass the fullHEDGEHOGpipeline\. Dragonfly is the best performing model in this KRAS G12D benchmark instance, retaining nearly twice as many final molecules as the second\-best REINVENT4 \(V\) model with 182 molecules\. The result is consistent with the design of Dragonfly, which combines interactome\-based learning with molecular generative priors and can incorporate ligand templates, 3D binding site information, and physicochemical property constraints\. The result suggests that protein conditioning is most effective when it is combined with strong chemical priors that preserve scaffold coherence, structural plausibility, and drug\-like molecular features\. The comparison between Dragonfly and Dragonfly \(b\) illustrates this point\. Dragonfly \(b\), which uses target\-ligand physicochemical descriptor generation, passes a descriptor stage similar to the out\-of\-the\-box Dragonfly model, retaining9,9909,990molecules and9,9679,967, respectively\. However, it falls to8787molecules after structural alerts evaluation, while the out\-of\-the\-box Dragonfly retains2,4882,488molecules\. This indicates that matching ligand descriptor profiles can improve early physicochemical compatibility, but does not by itself prevent structural liabilities\.

Across230,000230,000initial molecules, only1,4901,490molecules survive all stages\. The stage\-wise results show that benchmark success is highly local, because model rankings vary across different filtering stages\. The diagnostic value of HEDGEHOG is not limited to ranking unconditional, ligand\-based, and protein\-based models\. The stage\-wise failures also separate architectural and training objective effects\. Scaffold\- and motif\-aware graph models tend to preserve chemically coherent structures through early filters, ligand\-based models can improve target\-specific relevance while losing robustness under synthesis and structural alert constraints, and protein\-based 3D generators can show stronger pose\-level survival only after severe target\-agnostic attrition\. Thus, the main failure mode is not a single missing property but the inability to maintain chemical realism, synthesizability, and target compatibility simultaneously\. Results show thatHEDGEHOGis a stringent and highly discriminative benchmark that reduces the candidate pool under a realistic multi\-stage computational triage workflow and clearly separates performance across model classes\.

## 5Conclusions

Molecular generator evaluation is misleading when based on a limited set of scores and metrics\. Our results suggest that generator evaluation should report end\-to\-end survival across medicinal chemistry, structure\-based, and synthesis filters\. We show that evaluated generators are still poorly aligned with stringent end\-to\-end computational triage\. UnderHEDGEHOG, the large majority of generated molecules fail before the proxy stages, and success at one isolated stage does not guarantee survival through the full pipeline\. Under this KRAS G12D benchmark instance, Dragonfly has the highest end\-to\-end survival with345345final molecules, followed by REINVENT4 \(V\) with182182, and REINVENT4 with163163final molecules\. The benchmark therefore exposes a central limitation of the current paradigm, that many models can generate molecules that look promising under partial or local criteria, yet very few produce outputs that remain acceptable under a broader sequence of medicinal\-chemistry\-oriented constraints\.

A limitation of our evaluation is that we do not estimate variance across independent generation seeds for all models\. Some baseline implementations did not expose or consistently document seed control, and for fairness we therefore use a common evaluation protocol based on10,00010,000generated SMILES strings per model\. This benchmark instance is centered on the KRAS G12D switch\-II pocket and includes a target\-specific Asp12 interaction requirement\. Model rankings may therefore change for other targets, pocket classes, or interaction hypotheses\.

While our pipeline integrates physicochemical, structural, synthesis, docking, and post\-docking 3D evaluation stages, it does not yet explicitly model long\-term ADMET properties, pharmacokinetics, or clinical\-success predictors\. Future work should extend the benchmark with uncertainty\-aware generative modeling, multi\-objective optimization under experimentally grounded constraints, broader target coverage, and prospective validation\. By shifting focus from isolated benchmark metrics toward survival in actionable chemical space with molecules that satisfy basic medicinal chemistry criteria,HEDGEHOGprovides a stricter and more practically motivated framework for evaluating molecular generators\.

## Data Availability statement

## References

- Alhossary et al\. \[2015\]Amr Alhossary, Stephanus Daniel Handoko, Yuguang Mu, and Chee\-Keong Kwoh\.Fast, accurate, and reliable molecular docking with quickvina 2\.*Bioinformatics*, 31\(13\):2214–2216, 2015\.
- Atz et al\. \[2024\]Kenneth Atz, Leandro Cotos, Clemens Isert, Maria Håkansson, Dorota Focht, Mattis Hilleke, David F Nippa, Michael Iff, Jann Ledergerber, Carl CG Schiebroek, et al\.Prospective de novo drug design with deep interactome learning\.*Nature Communications*, 15\(1\):3408, 2024\.
- Baell and Holloway \[2010\]Jonathan B Baell and Georgina A Holloway\.New substructure filters for removal of pan assay interference compounds \(pains\) from screening libraries and for their exclusion in bioassays\.*Journal of medicinal chemistry*, 53\(7\):2719–2740, 2010\.
- Bagal et al\. \[2021\]Viraj Bagal, Rishal Aggarwal, PK Vinod, and U Deva Priyakumar\.Molgpt: molecular generation using a transformer\-decoder model\.*Journal of chemical information and modeling*, 62\(9\):2064–2076, 2021\.
- Baillif et al\. \[2024\]Benoit Baillif, Jason Cole, Patrick McCabe, and Andreas Bender\.Benchmarking structure\-based three\-dimensional molecular generative models using genbench3d: ligand conformation quality matters\.*arXiv preprint arXiv:2407\.04424*, 2024\.
- Bedart et al\. \[2024\]Corentin Bedart, Conrad Veranso Simoben, and Matthieu Schapira\.Emerging structure\-based computational methods to screen the exploding accessible chemical space\.*Current Opinion in Structural Biology*, 86:102812, 2024\.
- Bickerton et al\. \[2012\]G Richard Bickerton, Gaia V Paolini, Jérémy Besnard, Sorel Muresan, and Andrew L Hopkins\.Quantifying the chemical beauty of drugs\.*Nature chemistry*, 4\(2\):90–98, 2012\.
- Bilodeau et al\. \[2022\]Camille Bilodeau, Wengong Jin, Tommi Jaakkola, Regina Barzilay, and Klavs F Jensen\.Generative models for molecular discovery: Recent advances and challenges\.*Wiley Interdisciplinary Reviews: Computational Molecular Science*, 12\(5\):e1608, 2022\.
- Bouysset and Fiorucci \[2021\]Cédric Bouysset and Sébastien Fiorucci\.Prolif: a library to encode molecular interactions as fingerprints\.*Journal of cheminformatics*, 13\(1\):72, 2021\.
- Brenk et al\. \[2008\]Ruth Brenk, Alessandro Schipani, Daniel James, Agata Krasowski, Ian Hugh Gilbert, Julie Frearson, and Paul Graham Wyatt\.Lessons learnt from assembling screening libraries for drug discovery for neglected diseases\.*ChemMedChem: Chemistry Enabling Drug Discovery*, 3\(3\):435–444, 2008\.
- Brown et al\. \[2019\]Nathan Brown, Marco Fiscato, Marwin HS Segler, and Alain C Vaucher\.Guacamol: benchmarking models for de novo molecular design\.*Journal of chemical information and modeling*, 59\(3\):1096–1108, 2019\.
- Bruns and Watson \[2012\]Robert F Bruns and Ian A Watson\.Rules for identifying potentially reactive or promiscuous compounds\.*Journal of medicinal chemistry*, 55\(22\):9763–9772, 2012\.
- Dahlin and Walters \[2014\]Jayme L Dahlin and Michael A Walters\.The essential roles of chemistry in high\-throughput screening triage\.*Future medicinal chemistry*, 6\(11\):1265–1290, 2014\.
- David et al\. \[2020\]Laurianne David, Amol Thakkar, Rocío Mercado, and Ola Engkvist\.Molecular representations in ai\-driven drug discovery: a review and practical guide\.*Journal of cheminformatics*, 12\(1\):56, 2020\.
- Du et al\. \[2024\]Yuanqi Du, Arian R Jamasb, Jeff Guo, Tianfan Fu, Charles Harris, Yingheng Wang, Chenru Duan, Pietro Liò, Philippe Schwaller, and Tom L Blundell\.Machine learning\-aided generative molecular design\.*Nature Machine Intelligence*, 6\(6\):589–604, 2024\.
- Duffy et al\. \[2012\]Bryan C Duffy, Lei Zhu, Hélène Decornez, and Douglas B Kitchen\.Early phase drug discovery: cheminformatics and computational techniques in identifying lead series\.*Bioorganic & medicinal chemistry*, 20\(18\):5324–5342, 2012\.
- Ertl and Schuffenhauer \[2009\]Peter Ertl and Ansgar Schuffenhauer\.Estimation of synthetic accessibility score of drug\-like molecules based on molecular complexity and fragment contributions\.*Journal of cheminformatics*, 1\(1\):8, 2009\.
- Fawcett \[1950\]Frank S Fawcett\.Bredt’s rule of double bonds in atomic\-bridged\-ring structures\.*Chemical Reviews*, 47\(2\):219–274, 1950\.
- Friesner et al\. \[2004\]Richard A Friesner, Jay L Banks, Robert B Murphy, Thomas A Halgren, Jasna J Klicic, Daniel T Mainz, Matthew P Repasky, Eric H Knoll, Mee Shelley, Jason K Perry, et al\.Glide: a new approach for rapid, accurate docking and scoring\. 1\. method and assessment of docking accuracy\.*Journal of medicinal chemistry*, 47\(7\):1739–1749, 2004\.
- Frolova et al\. \[2025\]Daria Frolova, Talgat Daulbaev, Egor Sevriugov, Sergei A Nikolenko, Dmitry N Ivankov, Ivan Oseledets, and Marina A Pak\.Matcha: Multi\-stage riemannian flow matching for accurate and physically valid molecular docking\.*arXiv preprint arXiv:2510\.14586*, 2025\.
- García\-Ortegón et al\. \[2022\]Miguel García\-Ortegón, Gregor NC Simm, Austin J Tripp, José Miguel Hernández\-Lobato, Andreas Bender, and Sergio Bacallado\.Dockstring: easy molecular docking yields better benchmarks for ligand design\.*Journal of chemical information and modeling*, 62\(15\):3486–3502, 2022\.
- Gaulton et al\. \[2012\]Anna Gaulton, Louisa J Bellis, A Patricia Bento, Jon Chambers, Mark Davies, Anne Hersey, Yvonne Light, Shaun McGlinchey, David Michalovich, Bissan Al\-Lazikani, et al\.Chembl: a large\-scale bioactivity database for drug discovery\.*Nucleic acids research*, 40\(D1\):D1100–D1107, 2012\.
- Genheden et al\. \[2020\]Samuel Genheden, Amol Thakkar, Veronika Chadimová, Jean\-Louis Reymond, Ola Engkvist, and Esben Bjerrum\.Aizynthfinder: a fast, robust and flexible open\-source software for retrosynthetic planning\.*Journal of cheminformatics*, 12\(1\):70, 2020\.
- Ghazi Vakili et al\. \[2025\]Mohammad Ghazi Vakili, Christoph Gorgulla, Jamie Snider, AkshatKumar Nigam, Dmitry Bezrukov, Daniel Varoli, Alex Aliper, Daniil Polykovsky, Krishna M Padmanabha Das, Huel Cox Iii, et al\.Quantum\-computing\-enhanced algorithm unveils potential kras inhibitors\.*Nature Biotechnology*, pages 1–6, 2025\.
- Ghose et al\. \[1999\]Arup K Ghose, Vellarkad N Viswanadhan, and John J Wendoloski\.A knowledge\-based approach in designing combinatorial or medicinal chemistry libraries for drug discovery\. 1\. a qualitative and quantitative characterization of known drug databases\.*Journal of combinatorial chemistry*, 1\(1\):55–68, 1999\.
- Gong et al\. \[2024\]Haisong Gong, Qiang Liu, Shu Wu, and Liang Wang\.Text\-guided molecule generation with diffusion language model\.*Proceedings of the AAAI Conference on Artificial Intelligence*, 38:109–117, 2024\.
- Guan et al\. \[2023\]Jiaqi Guan, Wesley Wei Qian, Xingang Peng, Yufeng Su, Jian Peng, and Jianzhu Ma\.3d equivariant diffusion for target\-aware molecule generation and affinity prediction\.*arXiv preprint arXiv:2303\.03543*, 2023\.
- Hann et al\. \[1999\]Mike Hann, Brian Hudson, Xiao Lewell, Rob Lifely, Luke Miller, and Nigel Ramsden\.Strategic pooling of compounds for high\-throughput screening\.*Journal of chemical information and computer sciences*, 39\(5\):897–902, 1999\.
- Ho et al\. \[2020\]Jonathan Ho, Ajay Jain, and Pieter Abbeel\.Denoising diffusion probabilistic models\.*Advances in neural information processing systems*, 33:6840–6851, 2020\.
- Hoogeboom et al\. \[2022\]Emiel Hoogeboom, Vıctor Garcia Satorras, Clément Vignac, and Max Welling\.Equivariant diffusion for molecule generation in 3d\.*Proceedings of the International Conference on Machine Learning*, pages 8867–8887, 2022\.
- Huang et al\. \[2021\]Kexin Huang, Tianfan Fu, Wenhao Gao, Yue Zhao, Yusuf Roohani, Jure Leskovec, Connor W Coley, Cao Xiao, Jimeng Sun, and Marinka Zitnik\.Therapeutics data commons: Machine learning datasets and tasks for drug discovery and development\.*arXiv preprint arXiv:2102\.09548*, 2021\.
- Irwin and Shoichet \[2005\]John J Irwin and Brian K Shoichet\.Zinc\- a free database of commercially available compounds for virtual screening\.*Journal of chemical information and modeling*, 45\(1\):177–182, 2005\.
- Ivanenkov et al\. \[2023\]Yan Ivanenkov, Bogdan Zagribelnyy, Alex Malyshev, Sergei Evteev, Victor Terentiev, Petrina Kamya, Dmitry Bezrukov, Alex Aliper, Feng Ren, and Alex Zhavoronkov\.The hitchhiker’s guide to deep learning driven generative chemistry\.*ACS Medicinal Chemistry Letters*, 14\(7\):901–915, 2023\.
- Ivanenkov et al\. \[2019\]Yan A Ivanenkov, Bogdan A Zagribelnyy, and Vladimir A Aladinskiy\.Are we opening the door to a new era of medicinal chemistry or being collapsed to a chemical singularity? perspective\.*Journal of medicinal chemistry*, 62\(22\):10026–10043, 2019\.
- Jain \[2003\]Ajay N Jain\.Surflex: fully automatic flexible molecular docking using a molecular similarity\-based search engine\.*Journal of medicinal chemistry*, 46\(4\):499–511, 2003\.
- Jin et al\. \[2018\]Wengong Jin, Regina Barzilay, and Tommi Jaakkola\.Junction tree variational autoencoder for molecular graph generation\.*Proceedings of the International Conference on Machine Learning*, pages 2323–2332, 2018\.
- Jin et al\. \[2020\]Wengong Jin, Regina Barzilay, and Tommi Jaakkola\.Hierarchical generation of molecular graphs using structural motifs\.*Proceedings of the International Conference on Machine Learning*, pages 4839–4848, 2020\.
- Kessler et al\. \[2019\]Dirk Kessler, Michael Gmachl, Andreas Mantoulidis, Laetitia J Martin, Andreas Zoephel, Moriz Mayer, Andreas Gollner, David Covini, Silke Fischer, Thomas Gerstberger, et al\.Drugging an undruggable pocket on kras\.*Proceedings of the National Academy of Sciences*, 116\(32\):15823–15829, 2019\.
- Kim et al\. \[2023\]Sunghwan Kim, Jie Chen, Tiejun Cheng, Asta Gindulyte, Jia He, Siqian He, Qingliang Li, Benjamin A Shoemaker, Paul A Thiessen, Bo Yu, et al\.Pubchem 2023 update\.*Nucleic acids research*, 51\(D1\):D1373–D1380, 2023\.
- Kingma and Welling \[2013\]Diederik P Kingma and Max Welling\.Auto\-encoding variational bayes\.*arXiv preprint arXiv:1312\.6114*, 2013\.
- Koes et al\. \[2013\]David Ryan Koes, Matthew P Baumgartner, and Carlos J Camacho\.Lessons learned in empirical scoring with smina from the csar 2011 benchmarking exercise\.*Journal of chemical information and modeling*, 53\(8\):1893–1904, 2013\.
- Kwon and Lee \[2021\]Yongbeom Kwon and Juyong Lee\.Molfinder: an evolutionary algorithm for the global optimization of molecular properties and the extensive exploration of chemical space using smiles\.*Journal of cheminformatics*, 13\(1\):24, 2021\.
- Lagorce et al\. \[2017\]David Lagorce, Lina Bouslama, Jerome Becot, Maria A Miteva, and Bruno O Villoutreix\.Faf\-drugs4: free adme\-tox filtering computations for chemical biology and early stages drug discovery\.*Bioinformatics*, 33\(22\):3658–3660, 2017\.
- Landrum \[2013\]Greg Landrum\.Rdkit documentation\.*Release*, 1\(1\-79\):4, 2013\.
- Lipinski \[2004\]Christopher A Lipinski\.Lead\-and drug\-like compounds: the rule\-of\-five revolution\.*Drug discovery today: Technologies*, 1\(4\):337–341, 2004\.
- Lipinski et al\. \[1997\]Christopher A Lipinski, Franco Lombardo, Beryl W Dominy, and Paul J Feeney\.Experimental and computational approaches to estimate solubility and permeability in drug discovery and development settings\.*Advanced drug delivery reviews*, 23\(1\-3\):3–25, 1997\.
- Liu et al\. \[2024\]Songtao Liu, Zhengkai Tu, Hanjun Dai, and Peng Liu\.Sddbench: A benchmark for synthesizable drug design\.*arXiv preprint arXiv:2409\.05822*, 2024\.
- Loeffler et al\. \[2024\]Hannes H Loeffler, Jiazhen He, Alessandro Tibo, Jon Paul Janet, Alexey Voronov, Lewis H Mervin, and Ola Engkvist\.Reinvent 4: modern ai–driven generative molecule design\.*Journal of Cheminformatics*, 16\(1\):20, 2024\.
- Löffler \[2025\]Hannes Löffler\.Reinvent4 priors\.2025\.doi:10\.5281/zenodo\.15641297\.URL[https://doi\.org/10\.5281/zenodo\.15641297](https://doi.org/10.5281/zenodo.15641297)\.
- Mao et al\. \[2022\]Zhongwei Mao, Hongying Xiao, Panpan Shen, Yu Yang, Jing Xue, Yunyun Yang, Yanguo Shang, Lilan Zhang, Xin Li, Yuying Zhang, et al\.Kras \(g12d\) can be targeted by potent inhibitors via formation of salt bridge\.*Cell discovery*, 8\(1\):5, 2022\.
- Mary et al\. \[2024\]Hadrien Mary, Emmanuel Noutahi, Dom Invivo, Michel Moreau, Steven Pak, Desmond Gilmour, Shawn Whitfield, Valence\-Jonny Hsu, Honoré Hounwanou, Ishan Kumar, Saurav Maheshkar, Shuya Nakata, Kyle M\. Kovary, Cas Wognum, Michael Craig, and Deep Source Bot\.datamol\-io/datamol: 0\.12\.3, January 2024\.URL[https://doi\.org/10\.5281/zenodo\.10535844](https://doi.org/10.5281/zenodo.10535844)\.
- Maziarz et al\. \[2021\]Krzysztof Maziarz, Henry Jackson\-Flux, Pashmina Cameron, Finton Sirockin, Nadine Schneider, Nikolaus Stiefl, Marwin Segler, and Marc Brockschmidt\.Learning to extend molecular scaffolds with structural motifs\.*arXiv preprint arXiv:2103\.03864*, 2021\.
- McNutt et al\. \[2021\]Andrew T McNutt, Paul Francoeur, Rishal Aggarwal, Tomohide Masuda, Rocco Meli, Matthew Ragoza, Jocelyn Sunseri, and David Ryan Koes\.Gnina 1\.0: molecular docking with deep learning\.*Journal of cheminformatics*, 13\(1\):43, 2021\.
- Mistryukova et al\. \[2025\]Lukia Mistryukova, Vladimir Manuilov, Konstantin Avchaciov, and Peter O Fedichev\.Protobind\-diff: A structure\-free diffusion language model for protein sequence\-conditioned ligand design\.*bioRxiv*, pages 2025–06, 2025\.
- Nie et al\. \[2024\]Dou Nie, Huifeng Zhao, Odin Zhang, Gaoqi Weng, Hui Zhang, Jieyu Jin, Haitao Lin, Yufei Huang, Liwei Liu, Dan Li, et al\.Durian: a comprehensive benchmark for structure\-based 3d molecular generation\.*Journal of Chemical Information and Modeling*, 65\(1\):173–186, 2024\.
- Nigam et al\. \[2023\]AkshatKumar Nigam, Robert Pollice, Gary Tom, Kjell Jorner, John Willes, Luca Thiede, Anshul Kundaje, and Alán Aspuru\-Guzik\.Tartarus: A benchmarking platform for realistic and practical inverse molecular design\.*Advances in Neural Information Processing Systems*, 36:3263–3306, 2023\.
- Nikolenko \[2026\]Sergei Nikolenko\.Posecheck fast\.[https://github\.com/LigandPro/posecheck\-fast](https://github.com/LigandPro/posecheck-fast), 2026\.
- Noutahi et al\. \[2025\]Emmanuel Noutahi, Hadrien Mary, Kyle M\. Kovary, Shawn Whitfield, Julien St\-Laurent, Honoré Hounwanou, and Michael Craig\.datamol\-io/medchem: Molecular filtering for drug discovery, January 2025\.URL[https://doi\.org/10\.5281/zenodo\.14588938](https://doi.org/10.5281/zenodo.14588938)\.
- Pak et al\. \[2025\]Marina A Pak, Daria Frolova, Sergei A Nikolenko, Talgat Daulbaev, Daria Ryabchenko, Anna Litvin, Pavel Gurevich, Kamil Garifullin, Alexander Shapeev, Ivan Oseledets, et al\.Bento: Benchmarking classical and ai docking on drug design\-relevant data\.*bioRxiv*, pages 2025–12, 2025\.
- Passaro et al\. \[2025\]Saro Passaro, Gabriele Corso, Jeremy Wohlwend, Mateo Reveiz, Stephan Thaler, Vignesh Ram Somnath, Noah Getz, Tally Portnoi, Julien Roy, Hannes Stark, et al\.Boltz\-2: Towards accurate and efficient binding affinity prediction\.*BioRxiv*, pages 2025–06, 2025\.
- Pearce et al\. \[2006\]Bradley C Pearce, Michael J Sofia, Andrew C Good, Dieter M Drexler, and David A Stock\.An empirical process for the design of high\-throughput screening deck filters\.*Journal of chemical information and modeling*, 46\(3\):1060–1068, 2006\.
- Peng et al\. \[2022\]Xingang Peng, Shitong Luo, Jiaqi Guan, Qi Xie, Jian Peng, and Jianzhu Ma\.Pocket2mol: Efficient molecular sampling based on 3d protein pockets\.*Proceedings of the International Conference on Machine Learning*, pages 17644–17655, 2022\.
- Polykovskiy et al\. \[2020\]Daniil Polykovskiy, Alexander Zhebrak, Benjamin Sanchez\-Lengeling, Sergey Golovanov, Oktai Tatanov, Stanislav Belyaev, Rauf Kurbanov, Aleksey Artamonov, Vladimir Aladinskiy, Mark Veselov, et al\.Molecular sets \(moses\): a benchmarking platform for molecular generation models\.*Frontiers in pharmacology*, 11:565644, 2020\.
- Preuer et al\. \[2018\]Kristina Preuer, Philipp Renz, Thomas Unterthiner, Sepp Hochreiter, and Gunter Klambauer\.Fréchet chemnet distance: a metric for generative models for molecules in drug discovery\.*Journal of chemical information and modeling*, 58\(9\):1736–1741, 2018\.
- Schneuing et al\. \[2024\]Arne Schneuing, Charles Harris, Yuanqi Du, Kieran Didi, Arian Jamasb, Ilia Igashov, Weitao Du, Carla Gomes, Tom L Blundell, Pietro Lio, et al\.Structure\-based drug design with equivariant diffusion models\.*Nature Computational Science*, 4\(12\):899–909, 2024\.
- Schneuing et al\. \[2025\]Arne Schneuing, Ilia Igashov, Adrian W Dobbelstein, Thomas Castiglione, Michael Bronstein, and Bruno Correia\.Multi\-domain distribution learning for de novo drug design\.*arXiv preprint arXiv:2508\.17815*, 2025\.
- Schuffenhauer et al\. \[2020\]Ansgar Schuffenhauer, Nadine Schneider, Samuel Hintermann, Douglas Auld, Jutta Blank, Simona Cotesta, Caroline Engeloch, Nikolas Fechner, Christoph Gaul, Jerome Giovannoni, et al\.Evolution of novartis’ small molecule screening deck design\.*Journal of medicinal chemistry*, 63\(23\):14425–14447, 2020\.
- Segler et al\. \[2018\]Marwin HS Segler, Mike Preuss, and Mark P Waller\.Planning chemical syntheses with deep neural networks and symbolic ai\.*Nature*, 555\(7698\):604–610, 2018\.
- Shen et al\. \[2021\]Panpan Shen, Yu Yang, Yunyun Yang, Lilan Zhang, Chun\-Chi Chen, and Ray\-Ting Guo\.Gdp\-bound kras g12d in complex with th\-z816, 2021\.URL[https://doi\.org/10\.2210/pdb7ew9/pdb](https://doi.org/10.2210/pdb7ew9/pdb)\.
- Sterling and Irwin \[2015\]Teague Sterling and John J Irwin\.Zinc 15–ligand discovery for everyone\.*Journal of chemical information and modeling*, 55\(11\):2324–2337, 2015\.
- Swanson et al\. \[2024\]Kyle Swanson, Gary Liu, Denise B Catacutan, Autumn Arnold, James Zou, and Jonathan M Stokes\.Generative ai for designing and validating easily synthesizable and structurally novel antibiotics\.*Nature machine intelligence*, 6\(3\):338–353, 2024\.
- Thakkar et al\. \[2021\]Amol Thakkar, Veronika Chadimová, Esben Jannik Bjerrum, Ola Engkvist, and Jean\-Louis Reymond\.Retrosynthetic accessibility score \(rascore\)–rapid machine learned synthesizability classification from ai driven retrosynthetic planning\.*Chemical science*, 12\(9\):3339–3349, 2021\.
- Thomas et al\. \[2024\]Morgan Thomas, Noel M O’Boyle, Andreas Bender, and Chris De Graaf\.Molscore: a scoring, evaluation and benchmarking framework for generative models in de novo drug design\.*Journal of Cheminformatics*, 16\(1\):64, 2024\.
- Tropsha et al\. \[2024\]Alexander Tropsha, Olexandr Isayev, Alexandre Varnek, Gisbert Schneider, and Artem Cherkasov\.Integrating qsar modelling and deep learning in drug discovery: the emergence of deep qsar\.*Nature Reviews Drug Discovery*, 23\(2\):141–155, 2024\.
- Trott and Olson \[2010\]Oleg Trott and Arthur J Olson\.Autodock vina: improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading\.*Journal of computational chemistry*, 31\(2\):455–461, 2010\.
- Vasta et al\. \[2022\]James D Vasta, D Matthew Peacock, Qinheng Zheng, Joel A Walker, Ziyang Zhang, Chad A Zimprich, Morgan R Thomas, Michael T Beck, Brock F Binkowski, Cesear R Corona, et al\.Kras is vulnerable to reversible switch\-ii pocket engagement in cells\.*Nature chemical biology*, 18\(6\):596–604, 2022\.
- Veber et al\. \[2002\]Daniel F Veber, Stephen R Johnson, Hung\-Yuan Cheng, Brian R Smith, Keith W Ward, and Kenneth D Kopple\.Molecular properties that influence the oral bioavailability of drug candidates\.*Journal of medicinal chemistry*, 45\(12\):2615–2623, 2002\.
- Verdonk et al\. \[2003\]Marcel L Verdonk, Jason C Cole, Michael J Hartshorn, Christopher W Murray, and Richard D Taylor\.Improved protein–ligand docking using gold\.*Proteins: Structure, Function, and Bioinformatics*, 52\(4\):609–623, 2003\.
- Voršilák et al\. \[2020\]Milan Voršilák, Michal Kolář, Ivan Čmelo, and Daniel Svozil\.Syba: Bayesian estimation of synthetic accessibility of organic compounds\.*Journal of cheminformatics*, 12\(1\):35, 2020\.
- Walters and Namchuk \[2003\]W Patrick Walters and Mark Namchuk\.Designing screens: how to make your hits a hit\.*Nature reviews Drug discovery*, 2\(4\):259–266, 2003\.
- Weininger \[1988\]David Weininger\.Smiles, a chemical language and information system\. 1\. introduction to methodology and encoding rules\.*Journal of chemical information and computer sciences*, 28\(1\):31–36, 1988\.
- Xu and Stevenson \[2000\]Jun Xu and James Stevenson\.Drug\-like index: a new approach to measure drug\-like compounds and their diversity\.*Journal of Chemical Information and Computer Sciences*, 40\(5\):1177–1187, 2000\.
- Zhang et al\. \[2023\]Odin Zhang, Jintu Zhang, Jieyu Jin, Xujun Zhang, RenLing Hu, Chao Shen, Hanqun Cao, Hongyan Du, Yu Kang, Yafeng Deng, et al\.Resgen is a pocket\-aware 3d molecular generation model based on parallel multiscale modelling\.*Nature Machine Intelligence*, 5\(9\):1020–1030, 2023\.
- Zhao \[2024\]Hongtao Zhao\.The science and art of structure\-based virtual screening\.*ACS Medicinal Chemistry Letters*, 15\(4\):436–440, 2024\.
- Zhavoronkov et al\. \[2019\]Alex Zhavoronkov, Yan A Ivanenkov, Alex Aliper, Mark S Veselov, Vladimir A Aladinskiy, Anastasiya V Aladinskaya, Victor A Terentiev, Daniil A Polykovskiy, Maksim D Kuznetsov, Arip Asadulaev, et al\.Deep learning enables rapid identification of potent ddr1 kinase inhibitors\.*Nature biotechnology*, 37\(9\):1038–1040, 2019\.
- Zhou et al\. \[2024\]Guangfeng Zhou, Domnita\-Valeria Rusnac, Hahnbeom Park, Daniele Canzani, Hai Minh Nguyen, Lance Stewart, Matthew F Bush, Phuong Tran Nguyen, Heike Wulff, Vladimir Yarov\-Yarovoy, et al\.An artificial intelligence accelerated virtual screening platform for drug discovery\.*Nature Communications*, 15\(1\):7761, 2024\.
- Zhu et al\. \[2023\]Huimin Zhu, Renyi Zhou, Dongsheng Cao, Jing Tang, and Min Li\.A pharmacophore\-guided deep learning approach for bioactive molecular generation\.*Nature Communications*, 14\(1\):6234, 2023\.
- Zou et al\. \[2025\]Yurong Zou, Tao Guo, Zhiyuan Fu, Zhongning Guo, Weichen Bo, Dengjie Yan, Qiantao Wang, Jun Zeng, Dingguo Xu, Taijin Wang, et al\.A structure\-based framework for selective inhibitor design and optimization\.*Communications Biology*, 8\(1\):422, 2025\.

## Appendix

## ADetailed molecular workflow of theHEDGEHOGpipeline

The code is available at[https://github\.com/LigandPro/hedgehog](https://github.com/LigandPro/hedgehog)\. Unless otherwise stated, all experiments use the core benchmark dependencies listed in Table[4](https://arxiv.org/html/2607.13155#A1.T4)\.

Table 4:CoreHEDGEHOGbenchmark dependencies\. Main domain\-specific packages and tools used in the evaluation pipelineSupplementary File 1 reports the number of generated molecules that satisfy the corresponding threshold for eachHEDGEHOGfiltration stage, split by model class: unconditional, ligand\-based, and protein\-based\. The bold rows report the number of molecules that satisfy all filter thresholds simultaneously at each stage\.

At each stage, a molecule is evaluated at stages\+1s\+1only if it passed all criteria at stagess\. All thresholds were fixed before computing model\-level pass rates and were applied identically to all generated molecules and model classes\.

### A\.1Molecular preprocessing definitions

Stage 1 standardizes all generated SMILES strings so that downstream calculations are performed on a consistent and comparable set of molecules\. Different model architectures, training datasets and objectives, and sampling procedures can produce outputs in different formats\. For that reason,HEDGEHOGfirst parses each SMILES with DataMol and RDKit, repairs valence issues when possible, and sanitizes the resulting molecule\. Then explicit hydrogens are added to each molecule and checked for atom\-type compatibility with a predefined restricted set of1010elements: H, C, N, O, F, P, S, Cl, Br, I\. Although some approved drugs contain metals, boron, silicon, or other elements, this restricted atom set is appropriate for a general small\-molecule generation benchmark because it avoids model\-dependent failures caused by uncommon chemistry\.

Generated structures, especially from diffusion\-based models, may contain disconnected fragments, salts, or solvent molecules\. Therefore,HEDGEHOGretains only the largest molecular fragment and standardizes it by disconnecting metals, normalizing functional groups, and reionizing the molecule\. Salts and solvents are removed because they also may be artifacts of crystallization, and we evaluate the potential of a generator to produce a drug\-like molecule, rather than a full crystallographic assembly\. Finally, stereochemical annotations are removed and the molecule is converted to a canonical non\-isomeric SMILES representation\. This makes deduplication deterministic because standardized molecular strings can then be compared directly regardless of artifacts\. Figure[3](https://arxiv.org/html/2607.13155#A1.F3)summarizes per filter pass rates\.

![Refer to caption](https://arxiv.org/html/2607.13155v1/img/heatmaps/7.png)

\\phantomsubcaption

a![Refer to caption](https://arxiv.org/html/2607.13155v1/img/heatmaps/1.png)

\\phantomsubcaption

b![Refer to caption](https://arxiv.org/html/2607.13155v1/img/heatmaps/2.png)

\\phantomsubcaption

c![Refer to caption](https://arxiv.org/html/2607.13155v1/img/heatmaps/3.png)

\\phantomsubcaption

d![Refer to caption](https://arxiv.org/html/2607.13155v1/img/heatmaps/4.png)

\\phantomsubcaption

e![Refer to caption](https://arxiv.org/html/2607.13155v1/img/heatmaps/5.png)

\\phantomsubcaption

f![Refer to caption](https://arxiv.org/html/2607.13155v1/img/heatmaps/6.png)

Figure 3:HEDGEHOGmolecule counts for each stage\.a,Preprocessing\.b,Descriptors\.c,Structural filters\.d,Synthesis feasibility\.e,Docking and binding affinity\.f,Three\-dimensional filters\. Columns correspond to individual models grouped by model class\. Rows correspond to stage filters\. Color indicates the number of molecules passing each filter\. The Stage pass rows indicate the number of molecules passing all stage filters simultaneously\.
### A\.2Physicochemical descriptors definitions and thresholds

HEDGEHOGcomputes2121physicochemical descriptors using RDKit, including one custom molecular complexity term, MCE\-18\[Ivanenkov et al\.,[2019](https://arxiv.org/html/2607.13155#bib.bib34)\]\. All descriptors are two\-dimensional \(2D\) and do not depend on three\-dimensional \(3D\) conformations\.

No single literature rule set defines thresholds for all physicochemical descriptors used in theHEDGEHOGpanel\. Many medicinal chemistry filters were designed for specific screening pipelines and cannot be transferred directly to every generative model setting\. We therefore assigned thresholds in three ways, summarized in Table[5](https://arxiv.org/html/2607.13155#A1.T5)\. When a descriptor and cutoff matched a published filter directly, we used the reported threshold without modification\. For descriptors that were related to published rules but not defined in exactly the same way, we used the rule as a guide for the correspondingHEDGEHOGdescriptor\. Finally, for descriptors that were not covered by broad drug\-likeness filters, thresholds were chosen empirically after examining molecules\.

Table 5:Descriptor thresholds used in the physicochemical filtration panel\. The basis column indicates whether a threshold was taken directly from a published rule, adapted from a related literature filter, or set manually when no broadly accepted cutoff was availableFigure[4](https://arxiv.org/html/2607.13155#A1.F4)shows how generated molecules populate each descriptor range, whereas Figure[3](https://arxiv.org/html/2607.13155#A1.F3)summarizes how many molecules from each model satisfy each descriptor threshold\.

![Refer to caption](https://arxiv.org/html/2607.13155v1/img/total_descr_distrib.png)Figure 4:Physicochemical descriptor distributions for generated molecules across model classes\. Vertical dashed lines indicate lower \(red\) and upper \(blue\) thresholds\. Gray shaded regions correspond to values outside the accepted threshold range\.
### A\.3Structural filters definitions and thresholds

Structural filters remove molecules with structural motifs that are likely to be reactive, unstable, synthetically undesirable, or associated with well\-known medicinal chemistry liabilities\. Figure[3](https://arxiv.org/html/2607.13155#A1.F3)highlights pass rates of molecules during structural filters evaluation, and Table[6](https://arxiv.org/html/2607.13155#A1.T6)summarizes the structural filter configuration used in theHEDGEHOGpipeline, resulting in the set of2,4582,458underlying SMARTS patterns provided in Supplementary File 2\.

Table 6:Structural filters used in theHEDGEHOGpipeline
### A\.4Synthesis feasibility definitions and thresholds

We evaluate synthesis feasibility using four complementary metrics\. The Synthetic Accessibility score \(SA score\)\[Ertl and Schuffenhauer,[2009](https://arxiv.org/html/2607.13155#bib.bib17)\]is a fragment\- and complexity\-based heuristic that ranges molecules from 1, for molecules expected to be easy to synthesize, to 10, indicating molecules that are difficult to synthesize\. We apply an SA score cutoff of 4\.5\. The Retrosynthetic Accessibility score \(RA score\)\[Thakkar et al\.,[2021](https://arxiv.org/html/2607.13155#bib.bib72)\]is a machine\-learned classifier that predicts whether an automated retrosynthesis planner is likely to identify a route to the target\. We apply the natural decision cutoff of 0\.5 for the predicted probability to find a synthesis route\. SYBA\[Voršilák et al\.,[2020](https://arxiv.org/html/2607.13155#bib.bib79)\]is a fragment\-based Bernoulli Naive Bayes classifier\. Positive values indicate molecules expected to be easier to synthesize, and we therefore apply a threshold of 0\.

AiZynthFinder is a machine\-learned tool for retrosynthesis route planning\. In this work, we used a ZINC\-derived stock of building blocks from the ZINC database\[Sterling and Irwin,[2015](https://arxiv.org/html/2607.13155#bib.bib70)\], with a maximum search depth of 5, and a time limit of 300 seconds per molecule\. If AiZynthFinder finds a route within these constraints, the search is considered successful\. An example route proposed under this setup is shown in Figure[5](https://arxiv.org/html/2607.13155#A1.F5)\. Failure to find a route does not imply that the molecule is intrinsically non\-synthesizable, because the result depends on the selected stock, policies, depth, and time budget\. Model\-wise retention after the individual criteria and the final Stage 4 pass is summarized in Figure[3](https://arxiv.org/html/2607.13155#A1.F3)\.

![Refer to caption](https://arxiv.org/html/2607.13155v1/img/consensus.png)Figure 5:Example retrosynthetic route proposed by AiZynthFinder under the search configuration used in this work\.
### A\.5Docking and binding affinity definitions and thresholds

Stage 5 is the target\-aware structure\-based evaluation for KRAS G12D at the switch\-II pocket\.HEDGEHOGsupports multiple docking approaches, includingsmina, GNINA, Matcha and binding affinity estimation tool Boltz\-2\.

For molecular docking withsmina, GNINA, and Matcha, we use a search box of side length10​Å10~\\AAcentered at\[−1\.258,−12\.520,11\.390\]\[\-1\.258,\-12\.520,11\.390\], with exhaustiveness set to eight, nine output modes, and an energy range of3​k​c​a​l/m​o​l3kcal/mol\. Docking poses were retained if their predicted docking score was lower than \-6\.5​k​c​a​l/m​o​l6\.5kcal/mol\. This threshold was chosen to remove clearly unfavorable poses while retaining molecules with potentially meaningful target interactions in the binding pocket\. For Boltz\-2, we used the continuous affinity output defined aslog10⁡\(I​C50/1​μ​M\)\\log\_\{10\}\(IC\_\{50\}/1\\mu M\)\. Lower values correspond to stronger predicted binding\. We applied a cutoff oflog10⁡\(100\)=2\\log\_\{10\}\(100\)=2, corresponding to a predictedI​C50≤100​μ​MIC\_\{50\}\\leq 100\\mu M\. Figure[3](https://arxiv.org/html/2607.13155#A1.F3)shows per filter independent pass rates with an overall pass rate of all criteria simultaneously\.

The tools used in this stage make different assumptions about protein\-ligand binding\.Sminais a classical empirical docking engine, whereas GNINA augments the same general docking framework with convolutional neural network scoring\. Matcha is a recent docking pipeline based on multi\-stage flow matching with explicit physical validity filtering and GNINA\-based final ranking\. Boltz\-2 is a biomolecular foundation model that jointly predicts complex structure and binding affinity\. We therefore do not interpret a favorable score from any single tool as sufficient evidence of binding\. Instead, we prioritize molecules that remain favorable across multiple tools, because agreement across these distinct modeling paradigms is less likely to reflect a tool\-specific false positive result and provides a more conservative basis for subsequent pose\-level validation\.

Some models appear to optimize docking scores at the expense of chemical plausibility, resulting in favorable docking scores but chemically poor molecules\. To illustrate this, we sampled three representative molecules\. These molecules passed Stage 1 only and were docked into the KRAS G12D switch\-II pocket under the same conditions as in the main pipeline\. We report average docking score as the unweighted arithmetic mean of the best docking scores returned bysmina, GNINA, and Matcha\. Figure[6](https://arxiv.org/html/2607.13155#A1.F6)shows that these molecules satisfy the docking scoring function while violating basic medicinal chemistry or synthesizability constraints\. This highlights that a model may appear strong under a docking\-only metric by generating high\-scoring but chemically implausible candidates\.

![Refer to caption](https://arxiv.org/html/2607.13155v1/img/hackedDockScores.png)Figure 6:Examples of docking\-score hacked molecules\. Average docking score is the mean of three docking tools, evaluated in theHEDGEHOGpipeline\.
### A\.6Three\-dimensional filtration definitions and thresholds

At this stage, the goal is no longer to assess whether a molecule can in principle be placed in the pocket, but whether the retained docked pose is geometrically plausible, conformationally accessible, and consistent with the target\-specific interaction hypothesis\. Figure[3](https://arxiv.org/html/2607.13155#A1.F3)summarizes per filter pass rates with an overall pass rate of all criteria simultaneously\.

The final post\-docking stage applies a set of complementary 3D pose filters\. First, docked poses are required to remain inside the predefined docking box, ensuring that the final ligand geometry stays within the intended binding region\. In the KRAS G12D setting, this region was defined around the crystallographic 05C ligand\. The filter checks final ligand coordinates atom by atom and rejects any pose with atoms outside the box\. Second, pose quality is evaluated using a PoseCheck Fast\[Nikolenko,[2026](https://arxiv.org/html/2607.13155#bib.bib57)\]geometry check to remove poses with severe intermolecular clashes, excessive volumetric overlap with the protein, large displacement from the pocket, or internal steric inconsistencies\. The mean speed of0\.05​m​s0\.05msper pose comes from keeping the test purely geometric and batchable\. Heavy\-atom coordinates are evaluated with vectorized PyTorch distance calculations\. Third, conformer deviation is evaluated with a symmetry\-corrected RMSD backend to reject poses that deviate too strongly from accessible ligand conformations\. Each docked geometry was compared with a molecule\-specific RDKit ETKDGv3 conformer ensemble\. We used the minimum symmetry\-corrected RMSD to this ensemble and rejected poses above the3\.0​Å3\.0~\\AAthreshold\. Finally, protein–ligand interaction fingerprints computed with ProLIF are used to assess target\-specific interaction requirements\. For the KRAS G12D setting, interactions with Asp12 are biologically motivated because salt\-bridge formation with the mutant Asp12 residue has been reported as a route to KRAS G12D selectivity\[Mao et al\.,[2022](https://arxiv.org/html/2607.13155#bib.bib50)\]\. Figure[7](https://arxiv.org/html/2607.13155#A1.F7)illustrates this interaction\. This requirement is not presented as a universal rule for all targets, but as part of the example studied here and is motivated by the target biology and binding site context\.

![Refer to caption](https://arxiv.org/html/2607.13155v1/img/asp12.png)Figure 7:3D structure of KRAS\-G12D in complex with inhibitor TH\-Z816 \(PDB ID: 7EW9\)\.a, KRAS\-G12D bound to GDP, magnesium ion \(green\), and inhibitor TH\-Z816 \(CCD ID: 05C, dark grey\) in the switch\-II pocket displayed as a surface\.b, Key interaction between Asp 12 of KRAS\-G12D and the piperazine moiety of the inhibitor is shown with pink dashed lines\. Side chains of residues within5​Å5~\\AAof the ligand are displayed\.
### A\.7Per\-stage runtime

The HEDGEHOG workflow was implemented as a staged, parallelized screening cascade\. Early, low\-cost filters removed unsuitable molecules before computationally intensive retrosynthesis, docking, and affinity\-prediction steps were applied\. This design substantially reduced the number of candidates entering the slowest stages and enabled the complete benchmark to be run within practical wall\-clock times\. Per\-stage wall\-clock times are reported in Table[7](https://arxiv.org/html/2607.13155#A1.T7)\.

Table 7:Per\-stage runtime of theHEDGEHOGpipeline

## BModels overview

We categorize molecular generators by generative strategy and architecture because both impose distinct inductive biases, including validity and grammar errors for strings, geometry handling for 3D models, pocket alignment for pocket\-based models\[David et al\.,[2020](https://arxiv.org/html/2607.13155#bib.bib14)\],\[Bilodeau et al\.,[2022](https://arxiv.org/html/2607.13155#bib.bib8)\]\. This taxonomy is summarized in Table[8](https://arxiv.org/html/2607.13155#A2.T8)\.

Table 8:Taxonomy of molecular generators considered in our benchmark, by architecture \(rows\) and generative strategy \(columns\)\.Genetic algorithm \(GA\)\.GA is a heuristic optimizer that evolves molecules via crossover and mutation operations\. MolFinder\[Kwon and Lee,[2021](https://arxiv.org/html/2607.13155#bib.bib42)\]applies conformational space annealing to SMILES\[Weininger,[1988](https://arxiv.org/html/2607.13155#bib.bib81)\]and requires no generative model pretraining\.

Variational autoencoder \(VAE\)\.VAE models learn a latent distribution over chemical space with an encoder–decoder pair optimized through the evidence lower bound\[Kingma and Welling,[2013](https://arxiv.org/html/2607.13155#bib.bib40)\]\. JT\-VAE\[Jin et al\.,[2018](https://arxiv.org/html/2607.13155#bib.bib36)\], HierGraphVAE\[Jin et al\.,[2020](https://arxiv.org/html/2607.13155#bib.bib37)\], and MoLeR\[Maziarz et al\.,[2021](https://arxiv.org/html/2607.13155#bib.bib52)\]operate on graphs with scaffold\-aware decoders\. GENTRL\[Zhavoronkov et al\.,[2019](https://arxiv.org/html/2607.13155#bib.bib85)\]is a VAE with a structured latent prior and reinforcement\-learning fine\-tuning\.

Autoregressive models\.String\-based autoregressive models factorize sequence likelihood as∏iP​\(ti∣t<i\)\\prod\_\{i\}P\(t\_\{i\}\\mid t\_\{<i\}\)\. MolGPT\[Bagal et al\.,[2021](https://arxiv.org/html/2607.13155#bib.bib4)\]is a decoder\-only Transformer\. REINVENT4\[Loeffler et al\.,[2024](https://arxiv.org/html/2607.13155#bib.bib48)\]fine\-tunes a SMILES prior into an agent via policy gradient with recurrent or transformer\-based prior\.

For structure\-based design, 3D autoregressive models condition on a pocketPPand learn the conditional likelihood of a moleculeMMaspθ​\(M\|P\)=∏t=1Tpθ​\(zt\|z<t,P\)p\_\{\\theta\(M\|P\)\}=\\prod\_\{t=1\}^\{T\}p\_\{\\theta\(z\_\{t\}\|z\_\{<t\},P\)\}, where each stepztz\_\{t\}adds atoms, bonds, or coordinates\. Pocket2Mol\[Peng et al\.,[2022](https://arxiv.org/html/2607.13155#bib.bib62)\]and ResGen\[Zhang et al\.,[2023](https://arxiv.org/html/2607.13155#bib.bib83)\]are autoregressive 3D generators that build molecules inside the pocket\. Dragonfly\[Atz et al\.,[2024](https://arxiv.org/html/2607.13155#bib.bib2)\]is an interactome\-based graph\-to\-sequence model that can incorporate ligand templates or 3D binding site information by encoding pocket geometry with SE\(3\) or E\(3\)\-equivariant encoders\.

Pharmacophore\-based models use conditionccas a set of 3D interaction features and geometry introducing latentzzwhich, viap​\(x\|c\)=∫pθ​\(x\|c,z\)​p​\(z\)​𝑑zp\\left\(x\|c\\right\)=\\int p\_\{\\theta\(x\|c,z\)\}p\\left\(z\\right\)dz, models the many\-to\-many relationship between pharmacophores and ligands\. PGMG\[Zhu et al\.,[2023](https://arxiv.org/html/2607.13155#bib.bib87)\]represents a pharmacophore as a fully connected graph\. GCPG\[Zou et al\.,[2025](https://arxiv.org/html/2607.13155#bib.bib88)\]is a transformer encoder–decoder whose hidden state is modulated by pharmacophore embeddings\.

Diffusion models\.Diffusion models learn to approximate the reverse process of a predefined forward noising process\[Ho et al\.,[2020](https://arxiv.org/html/2607.13155#bib.bib29)\]\. E\(3\)DM\[Hoogeboom et al\.,[2022](https://arxiv.org/html/2607.13155#bib.bib30)\]is an E\(3\)\-equivariant model that jointly denoises atom coordinates and types\. DiffSBDD\[Schneuing et al\.,[2024](https://arxiv.org/html/2607.13155#bib.bib65)\]is an SE\(3\)\-equivariant 3D\-conditional model\. TargetDiff\[Guan et al\.,[2023](https://arxiv.org/html/2607.13155#bib.bib27)\]conditions the diffusion process on a protein binding site\. TGM\-DLM\[Gong et al\.,[2024](https://arxiv.org/html/2607.13155#bib.bib26)\]is a text\-guided diffusion language model over SMILES token embeddings\. In our benchmark, it is used in a non\-target\-specific prompting setting and is therefore grouped with unconditional baselines\. ProtoBind\-Diff\[Mistryukova et al\.,[2025](https://arxiv.org/html/2607.13155#bib.bib54)\]is a structure\-free diffusion language model that takes a protein amino\-acid sequence as a condition\.

Flow matching\.Flow\-matching models learn continuous\-time velocity fields that transport a base distribution to the data distribution\. DrugFlow\[Schneuing et al\.,[2025](https://arxiv.org/html/2607.13155#bib.bib66)\]is a pocket\-conditioned ligand generation model with flow\-based sampling\.

## CKRAS G12D benchmark instance

The KRAS G12D receptor was prepared from PDB 7EW9 \(PDB ID: pdb\_00007ew9\[Shen et al\.,[2021](https://arxiv.org/html/2607.13155#bib.bib69)\]\)\. The crystallographic inhibitor TH\-Z816, GDP, Mg2\+, and crystallographic water molecules were removed before docking\. Residue numbering was kept consistent with the PDB file, with Asp12 used as the target\-specific interaction residue\. The prepared receptor was used forsmina, GNINA, Matcha, and Boltz\-2 calculations and Stage 6 post\-docking checks\.

We used KRAS G12D as the benchmark target for HEDGEHOG\. KRAS G12D is the most prevalent KRAS oncogenic mutation\[Mao et al\.,[2022](https://arxiv.org/html/2607.13155#bib.bib50)\], and for many years was considered undruggable\. Unlike KRAS G12C, it lacks the Cys12 required for covalent binding, but at the same time the negatively charged Asp12 can be targeted via salt\-bridge interaction, as demonstrated in the development of TH\-Z816\[Kessler et al\.,[2019](https://arxiv.org/html/2607.13155#bib.bib38)\]\. Given its challenging nature and clinical significance, we considered KRAS G12D a relevant and mechanistically informative benchmark target for the HEDGEHOG evaluation pipeline\[Mao et al\.,[2022](https://arxiv.org/html/2607.13155#bib.bib50)\]\.

For unconditional models, KRAS G12D is used as the target in docking, affinity estimation and post\-docking 3D checks\. For protein\-based models, it also provides the KRAS G12D structure used as target\-conditioning input\. For ligand\-based models, we used a set of KRAS G12D inhibitors provided by\[Ghazi Vakili et al\.,[2025](https://arxiv.org/html/2607.13155#bib.bib24)\]as conditioning data\.

The initial set contains 645 molecules\. We assessed validity using RDKit and identified 13 molecules with invalid structures due to valency issues\. These molecules were removed, leaving 632 unique and valid ligands\.

## DREINVENT4 configurations

We evaluated REINVENT4 in onede novoconfiguration and three ligand\-conditioned configurations\. REINVENT4 was used for this analysis because the same framework provides public priors,de novoSMILES generation, molecule\-based generation, and transfer\-learning workflows\. This allows prior choice to be varied while keeping the downstream HEDGEHOG evaluation fixed, isolating how prior choice and transfer\-learning within models affect downstream utility\. For sampling, we used beam search with temperature set to1\.01\.0and number of SMILES per input SMILES set to500500\. All REINVENT4 configurations were evaluated on one NVIDIA A100 40GB GPU\. Prior checkpoints are listed in\[Löffler,[2025](https://arxiv.org/html/2607.13155#bib.bib49)\]and are open\-source\.

The unconditional baseline was denoted as REINVENT4 \(checkpoint,reinvent\.prior\) in sampling mode\. No KRAS ligand SMILES were supplied to this run\.

REINVENT4 \(V\) \(Vanilla, checkpoint:reinvent\.prior\)\. Prior was trained on ChEMBL data\[Gaulton et al\.,[2012](https://arxiv.org/html/2607.13155#bib.bib22)\]\. We input632632KRAS G12D inhibitors to the model and obtained49,11949,119valid SMILES strings in88hours\. The REINVENT4 \(P\) \(provided prior, checkpoint:mol2mol\_medium\_similarity\.prior\) was trained with an explicit medium Tanimoto similarity objective\. Because the public Mol2Mol prior was trained on PubChem\-derived molecular pairs\[Kim et al\.,[2023](https://arxiv.org/html/2607.13155#bib.bib39)\], we removed KRAS seed ligands with exact PubChem matches before sampling to prevent data leakage\. This filtering reduced the seed set to 583 molecules\. The model outputs126,001126,001valid SMILES strings in88hours\. For REINVENT4 \(TL\) \(transfer\-learning\), we fine\-tuned a REINVENT4 prior \(reinvent\.prior\) on632632known KRAS G12D inhibitors with the Mol2Mol training pipeline\. We defined a similarity threshold of0\.70\.7, and fine\-tuned for1,0001,000epochs for7272hours\. For further evaluation, we selected the checkpoint with the lowest loss on the validation set\.

Similar Articles

DrugGen 2: A disease-aware language model for enhancing drug discovery

Hugging Face Daily Papers

DrugGen-2 fine-tunes GPT-2 using supervised learning and reinforcement learning (GRPO) to generate small molecules conditioned on both disease ontology and target protein sequences, achieving superior diversity and binding affinity for drug discovery.