Generating Developable 3D Molecules via Pocket-Conditioned Diffusion and Property-Aware Optimization

arXiv cs.LG Papers

Summary

This paper introduces a novel diffusion-based generative model for structure-based drug design that decouples pocket and ligand representation learning and incorporates multi-scale interaction signals and property-aware optimization to generate developable 3D molecules with improved binding affinity and ADMET properties.

arXiv:2607.12349v1 Announce Type: new Abstract: Drug discovery and development is time-consuming and resource-intensive, motivating computational approaches such as diffusion models for de novo drug design. Many such models follow the structure-based drug design (SBDD) paradigm, generating molecules to fit a target binding pocket. However, existing diffusion-based SBDD methods typically couple pocket and ligand representation learning, model interactions only at the atom level, and prioritize binding affinity over other developability properties. Here, we introduce conDitar-dev, a conditional diffusion-based SBDD framework for generating ligands with strong binding affinities and favorable ADMET properties. It consists of three modules: msPRL, a pretrained multi-scale pocket representation learning module; conDitar, a pocket-conditioned diffusion model guided by msPRL representations; and paOPT, a generation-time method for optimizing ligand developability. On a newly curated benchmark of human disease targets, conDitar outperforms state-of-the-art SBDD baselines, achieving an average binding score of -8.85 kcal/mol. Across five ADMET properties, conDitar-dev improves performance by up to 73% over conDitar. To further validate the abilities of conDitar-dev to generate developable molecules, we have applied it to two validated druggable targets: programmed death-ligand 1 (PD-L1) and colony-stimulating factor 1 receptor (CSF1R) proteins. Top-ranked generatively designed molecules and their analogs have been experimentally synthesized and biologically tested. Two molecules generated directly by conDitar-dev for PD-L1 exhibited SPR-derived $K_D$ values of 3.49 and 3.75 $\mu$M, respectively. Hit expansion based on conDitar-dev-designed molecules identified selective CSF1R inhibitors with IC$_{50}$ values as low as 200 nM, while also uncovering opportunities for drug repositioning.
Original Article
View Cached Full Text

Cached at: 07/15/26, 04:18 AM

# Generating Developable 3D Molecules via Pocket-Conditioned Diffusion and Property-Aware Optimization
Source: [https://arxiv.org/html/2607.12349](https://arxiv.org/html/2607.12349)
Ruoxi Gao1, Jiangweizhi Peng2, Ziqi Chen3, Frazier N\. Baker1, David C\. Kombo4, John L\. Kane, Jr\.4, Andrew A\. Scholte4, Yi Li4, Matthew J\. LaMarche4, Luigi I\. Iconaru5, Hans\-Peter Biemann5, Mingyi Hong6, Xia Ning1,7,8,9 ๐Ÿ–‚

## Introduction

Drug discovery and development is complex, time\-consuming, and resource\-intensive โ€“ new drugs typically take 12\-15 years to develop\[singh2023drug\]at costs of $378 million to $1\.76 billion\[sertkaya2024costs\]\. To accelerate this process and reduce costs, computational methods, particularly recent generative models \(e\.g\., diffusion\[hoogeboom22diff,guan2023decompdiff\], variational autoencoders\[jin18jtvae\], large language models\[Yu2024\]\), have been employed for*de novo*drug design\. Rather than conducting expensive searches over large libraries to identify potential binding molecules, these models directly generate molecules*in silico*based on the chemical knowledge learned from a vast amount of experimental data\. These models generally follow one of the two conventional drug design frameworks:\(1\)structure\-based drug design \(SBDD\), where molecules are designed to fit into the binding pocket structure of a given target; and\(2\)ligand\-based drug design \(LBDD\), where molecules are designed to resemble a known binding ligand\. Compared to LBDD, SBDD offers a more direct and mechanistic entry point to molecular discovery by explicitly exploiting pocket\-ligand interactions, thereby providing a principled starting point for subsequent ligand\-based modeling and optimization\. Recent development of generative models for SBDD\[chen2025generating,guan2023decompdiff,guan2023targetdiff\]features diffusion\[ho2020ddpm\], a process that gradually transforms noise into molecular structures, as the leading paradigm\. By learning the underlying distributions of chemically and structurally feasible molecules from data and leveraging the distributions to generate novel molecules, diffusion holds substantial promise to revolutionize*de novo*drug design by rapidly producing high\-quality, functionally relevant molecular candidates\.

Still, existing diffusion\-based models for SBDD leave substantial room for improvement\. Typically, diffusion\-based SBDD models tend to couple pocket and ligand representation learning, using a shared representation to reflect the complex pocket\-ligand interactions\. However, such coupling could amplify early generation errors, leading to progressively degraded representations during diffusionโ€™s iterative refinement process\. Instead, dedicated pocket representation learning may better capture binding pocket structures and properties, and thus, ligand generation can be conditioned on robust pocket features that favor strong binding\. Furthermore, many existing models do not yet fully leverage interaction signals across multiple scales โ€“ ranging from atom\-level to residue\-level โ€“ in modeling pocket\-ligand interactions\. Incorporating rich, multi\-scale interaction signals could allow models to better represent the biochemical environment of the binding site, capturing both local atomic interactions and broader residue\-level context that jointly govern ligand recognition and binding\. In addition, generative SBDD models typically prioritize the generation of ligands with optimized binding affinity, leaving other critical drug properties, such as Absorption, Distribution, Metabolism, Excretion, and Toxicity \(ADMET\), to be addressed in downstream lead optimization phase\. While this aligns with the conventional drug development process\[hughes2011principles\], it also underscores the opportunity for a more proactive strategy that considers a broader set of drug\-relevant properties at the stage of initial molecule generation, thereby increasing the likelihood of ultimate success\. These insights motivate the design of a new, decoupled, multi\-scale, and developability\-aware diffusion\-based framework for SBDD\.

Here, we introduce๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits, a conditional diffusion\-based SBDD framework capable of generating ligands with strong binding affinities and favorable ADMET properties\.๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsconsists of three modules:\(1\)a pretrainedmulti\-scalepocketrepresentationlearning module, referred to as๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\\mathop\{\\mathsf\{msPRL\}\}\\limits, which encodes the binding pocketโ€™s atomic and residue\-level composition and structure into expressive representations that will guide ligand generation;\(2\)a pocket\-conditioneddiffusion model that generates ligands for strongtarget binding, referred to as๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limits; and\(3\)a generation\-time,property\-awareoptimization mechanism for drug developability built upon๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limits, referred to as๐—‰๐–บ๐–ฎ๐–ฏ๐–ณ\\mathop\{\\mathsf\{paOPT\}\}\\limits\.๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsemploys apocket\-conditionedligandgenerator, referred to as๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limits, to predict atom types and structures of a potential binding ligand from the pocket representation learned from๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\\mathop\{\\mathsf\{msPRL\}\}\\limitsand pocket\-ligand interactions at both the atom and residue levels\. The๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitsโ€™s predicted atom positions and types are used to direct the diffusion process of๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitstowards the final generated ligands of high binding affinities\. Beyond binding affinity, the training\-free, plug\-and\-play๐—‰๐–บ๐–ฎ๐–ฏ๐–ณ\\mathop\{\\mathsf\{paOPT\}\}\\limitssteers ligand generation towards both high binding affinities and favorable developability in terms of ADMET profiles\. In summary,๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitspresents a new diffusion\-based paradigm for generating ligands with high binding affinity and favorable drug developability profiles, bridging binding\-driven design and developability optimization in real\-world drug discovery and development settings\. Fig\.[1](https://arxiv.org/html/2607.12349#Sx1.F1)presents the overview architecture of๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\.

To evaluate๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits, as well as other SBDD methods in the literature\[luo2021sbdd,peng22pocket2mol,guan2023targetdiff,guan2023decompdiff\], we carefully curate a new dataset, denoted as๐–ข๐–ฃ๐–ง\\mathop\{\\mathsf\{CDH\}\}\\limits, on top of the widely\-used benchmark dataset CrossDocked2020\[Francoeur2020\], denoted as๐–ข๐–ฃ\\mathop\{\\mathsf\{CD\}\}\\limits\. While well adopted,๐–ข๐–ฃ\\mathop\{\\mathsf\{CD\}\}\\limitsincludes both human and non\-human protein targets \(33% and 67%, respectively\)\. This poses risks to translational efficacy and safety, as even for conserved targets, sequence, structural, dynamics, and conformational differences between human and non\-human proteins can alter ligand binding and target engagement\[marshall2023poor\]\. To address this issue,๐–ข๐–ฃ๐–ง\\mathop\{\\mathsf\{CDH\}\}\\limitsis curated to consist exclusively of human targets, including human targets from๐–ข๐–ฃ\\mathop\{\\mathsf\{CD\}\}\\limits, the human orthologs of non\-human targets in๐–ข๐–ฃ\\mathop\{\\mathsf\{CD\}\}\\limits, and some additional, carefully\-selected, experimentally validated human targets from the Protein Database Bank \(PDB\)\[berman2000protein\]that are of high therapeutic interest in life\-threatening conditions \(e\.g\., cancers\)\. Over๐–ข๐–ฃ๐–ง\\mathop\{\\mathsf\{CDH\}\}\\limitsand๐–ข๐–ฃ\\mathop\{\\mathsf\{CD\}\}\\limits, we extensively compare๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsagainst seven state\-of\-the\-art baselines in generating high\-quality ligands\. Computational results demonstrate that๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsgenerates ligands with superior predicted binding affinity, achieving average scores ofโˆ’8\.85\-8\.85kcal/mol and a 7\.4% improvement over the best baseline on๐–ข๐–ฃ๐–ง\\mathop\{\\mathsf\{CDH\}\}\\limits, and๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsfurther improves ADMET properties on๐–ข๐–ฃ๐–ง\\mathop\{\\mathsf\{CDH\}\}\\limitswith an average improvement of up to 73% across five ADMET properties while maintaining comparable predicted binding affinities\.

To further demonstrate that๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitscan propose developable molecules, we conduct case studies on the programmed death\-ligand 1 \(PD\-L1\), and colony\-stimulating factor\-1 receptor \(CSF1R\) proteins, which are two therapeutically important and validated targets with available structures\. Top\-ranked๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated ligands and their analogs are synthesized and tested in biological assays\. Across the two targets, the selected molecules show promising biological activities\. In particular, two ligands generated by๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsfor PD\-L1 show SPR\-derived binding activities in the low micromolar range, with KD\{\}\_\{\\text\{D\}\}values below 3\.80ฮผ\\muM\. For CSF1R, two analogs derived from๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated ligands show strong kinase activities and high selectivity, with IC50values below 0\.31ฮผ\\muM and promiscuity hit rates of at most 2\.4%\. These biological testing results suggest that๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsis not limited to generating ligands with computationally predicted binding, but can also produce biologically active and developable molecules*in vitro*\.

In summary, we highlight the advantages of๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsas follows:

- โ€ข๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsunifies pocket\-conditioned generation and property optimization within a single framework, incorporating developability considerations directly into the initial drug design process\.
- โ€ข๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsproduces ligands with strong binding affinities leveraging decoupled pocket representation learning\. This decoupling allows๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\\mathop\{\\mathsf\{msPRL\}\}\\limitsto better capture binding pocket information, thus providing๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitswith robust pocket conditioning\.
- โ€ข๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\\mathop\{\\mathsf\{msPRL\}\}\\limitsand๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitsencode multi\-scale features of the binding pocket and pocket\-ligand interactions, allowing molecule generation to be informed by both local atomic details and broader residue\-level context that jointly govern ligand recognition and binding\.
- โ€ข๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsachieves favorable ADMET profiles by incorporating the training\-free๐—‰๐–บ๐–ฎ๐–ฏ๐–ณ\\mathop\{\\mathsf\{paOPT\}\}\\limitsto directly optimize those properties during the generation process of๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limits\. This allows external ADMET signals and pocket conditioning to jointly guide ligand generation\.
- โ€ขCase studies with extensive*in silico*analyses on PD\-L1 and CSF1R demonstrate that๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitscan generate developable molecules with binding modes similar to those of known binding ligands of these two targets\.
- โ€ขFifteen top\-ranked๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-designed molecules for PD\-L1 and CSF1R, and their derivatives and analogs, have been experimentally synthesized and tested\. Two hits have been identified with target inhibition at nanomolar concentrations, and seven hits have target inhibition at low micromolar concentrations\.

![Refer to caption](https://arxiv.org/html/2607.12349v1/x1.png)Fig\. 1:The overall schematic diagram of๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\.a,Pocket representation pretraining,๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\\mathop\{\\mathsf\{msPRL\}\}\\limits\. A pocket representation learning module is pretrained to encode protein pockets into multi\-scale atom\-level and residue\-level features\.b,๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsmodel training and inference\. Conditioned on these fixed pocket features,๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsgenerates 3D ligands for target pockets through a pocket\-conditioned diffusion model\.c,Generation\-time molecule optimization,๐—‰๐–บ๐–ฎ๐–ฏ๐–ณ\\mathop\{\\mathsf\{paOPT\}\}\\limits\. During inference,๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsoptimizes the original trajectory in a training\-free manner to steer generated molecules toward improved ADMET properties\.
## Related Work

##### Generative Models for SBDD

Structure\-based drug design \(SBDD\) has leveraged various generative modeling approaches to design ligands tailored to specific protein pockets\. Peng*et al\.*\[peng22pocket2mol\]constructed an encoder\-predictor framework๐–ฏ๐—ˆ๐–ผ๐—„๐–พ๐—๐Ÿค๐–ฌ๐—ˆ๐—…\\mathop\{\\mathsf\{Pocket2Mol\}\}\\limitsthat can generate ligands in an autoregressive way\. However,๐–ฏ๐—ˆ๐–ผ๐—„๐–พ๐—๐Ÿค๐–ฌ๐—ˆ๐—…\\mathop\{\\mathsf\{Pocket2Mol\}\}\\limitsrelies on the autoregressive sampling process, which tends to violate geometric constraints and produce ligands with limited interactions with the pocket, resulting in low binding affinity\. More recently, target\-aware diffusion models\[guan2023targetdiff,guan2023decompdiff,zhoudecompopt,gu2024aligning,huang2024protein\]for ligand generation have largely addressed these limitations, offering better geometric consistency and binding affinities\. Guan*et al\.*\[guan2023targetdiff\]pioneered this direction by introducing an E\(3\)\-equivariant conditional diffusion model\[fuchs2020se,satorras2021n,ho2020ddpm,hoogeboom2021argmax\]๐–ณ๐–บ๐—‹๐—€๐–พ๐—๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{TargetDiff\}\}\\limitsthat jointly generates atomic coordinates of ligands in 3D and types conditioned on protein pocket structures\. It achieves this through careful design of equivariant score networks and denoising processes that respect geometric symmetries, thereby establishing a strong baseline for diffusion\-based SBDD methods\.

To further improve conformational stability and molecular validity, Guan*et al\.*\[guan2023decompdiff\]proposed๐–ฃ๐–พ๐–ผ๐—ˆ๐—†๐—‰๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{DecompDiff\}\}\\limits\. Inspired by medicinal chemistry practice of scaffold\-and\-arm design\[schneider1999scaffold\],๐–ฃ๐–พ๐–ผ๐—ˆ๐—†๐—‰๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{DecompDiff\}\}\\limitsdecomposes ligand molecules into scaffolds and arms and learns separate diffusion priors for each\. This decomposed design facilitates more structured generation and enhances efficiency in exploring chemical space\. Building upon the idea of a decomposed prior, Zhou*et al\.*\[zhoudecompopt\]introduced๐–ฃ๐–พ๐–ผ๐—ˆ๐—†๐—‰๐–ฎ๐—‰๐—\\mathop\{\\mathsf\{DecompOpt\}\}\\limits, which decomposes pocket\-ligand interactions into local subpocket\-arm interactions and employs iterative optimization over the arms\. However,๐–ณ๐–บ๐—‹๐—€๐–พ๐—๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{TargetDiff\}\}\\limits,๐–ฃ๐–พ๐–ผ๐—ˆ๐—†๐—‰๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{DecompDiff\}\}\\limitsand๐–ฃ๐–พ๐–ผ๐—ˆ๐—†๐—‰๐–ฎ๐—‰๐—\\mathop\{\\mathsf\{DecompOpt\}\}\\limitsprimarily rely on geometric features of the pocket and ligand\-only priors without directly encoding protein\-ligand interaction signals into the diffusion dynamics\. To explicitly leverage such interactions, Huang*et al\.*\[huang2024protein\]proposed๐–จ๐–ฏ๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{IPDiff\}\}\\limits, which redesigns both forward and reverse processes to be informed by a learnable binding\-affinity signal\. This adaptation guides the generation toward ligands with improved binding properties\. Whereas๐–จ๐–ฏ๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{IPDiff\}\}\\limitssteers trajectories using a specific supervised affinity signal, Gu*et al\.*\[gu2024aligning\]introduced a more general framework๐– ๐–ซ๐–จ๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{ALIDiff\}\}\\limits, which controls generation by aligning the reverse process of a pretrained diffusion model with desired preferences via Diffusion\-DPO\[wallace2024diffusion\]adapted to molecular space\. Beyond diffusion\-based methods, recent work has also explored flow matching as an alternative generative paradigm for SBDD\. For example, Cremer*et al\.*\[cremer2026flowr\]introducedFLOWR, which combines continuous and categorical flow matching with equivariant optimal transport for ligand generation\.

Despite the different designs of current generative models, existing methods entangle pocket encoding with ligand denoising through a single end\-to\-end generation objective, which can blur interaction signals\. This gap motivates our design of decoupling the two: obtain a stable pretrained pocket summary first, then use it as a fixed context, enabling denoising to focus solely on pocket\-ligand interactions\. Moreover, prior works process pocket information only at the atom level, which may miss higher\-level structural patterns encoded by residues\. In contrast, our approach combines both atom\- and residue\-level pocket information to capture pocket characteristics and protein\-ligand interactions at multiple scales\. Furthermore, most prior works focus on improving binding affinity while ignoring ADMET profiles\. Our approach seeks to address this issue by explicitly incorporating ADMET\-aware optimization into the generation process\.

##### Alignment for Diffusion Models

There is a growing body of research on adapting pretrained generative diffusion models to specific downstream objectives\. These approaches can be broadly categorized into training\-based and training\-free methods, depending on whether the model parameters are updated\. In training\-based methods, Prabhudesai*et al\.*\[prabhudesai2023aligning\]proposed directly backpropagating through differentiable reward functions into diffusion model parameters\. While effective when such reward functions exist, this approach is prone to reward hacking in the absence of strong regularization\. Black*et al\.*\[blacktraining\]applied policy gradient methods such as PPO to finetune diffusion models\. Although more stable than direct backpropagation, reinforcement learning finetuning remains resource\- and time\-intensive, limiting its flexibility for diverse downstream tasks\. By contrast, training\-free methods operate in a plug\-and\-play fashion, requiring minimal or no parameter updates\. Song*et al\.*\[song2023loss\]introduced reward\-guided sampling, incorporating estimated reward gradients as guidance terms during generation\. Li*et al\.*\[li2024derivative\]proposed a value\-based decoding strategy that selects denoised samples according to estimated value functions\. More recently, Ma*et al\.*\[ma2025inference\]framed alignment as a search problem and introduced a search\-over\-paths algorithm to exploit generation\-time scaling behaviors\. While these methods advance training\-free alignment, they share a key limitation: it is difficult to reliably assess the quality of generated samples before the generative process is complete\.

To address the above\-mentioned limitations, a recent emerging line of work formulates the entire generation process as an optimization problem, where the objective is to directly optimize task\-specific downstream performance of the final generated sample\. Concretely, in standard diffusion sampling, each denoising step draws the next state from the model\-predicted posterior by combining a deterministic update \(the posterior mean\) with stochastic noise injected by the sampler\. By treating the injected noise variables as optimization variables, one can perturb the denoising trajectory and thereby steer the final sample toward preferred outcomes while keeping the diffusion model fixed\. Eyring*et al\.*\[eyring2024reno\]optimized the noise variable at the initial step of diffusion generation against external reward functions, achieving improved alignment at inference\. Tang*et al\.*\[tanginference\]extended this by optimizing injected noise across the entire denoising trajectory, propagating reward signals on the final outputs back to noise inputs while keeping the model parameters fixed\. While technically sound and effective in the field of image generation, these methods could not be readily applied for generating new chemical entities \(NCEs\)\. In particular, they are typically designed for continuous Gaussian noise and differentiable objectives, making it nontrivial to handle discrete \(categorical\) distribution over atom types and to incorporate molecular property evaluators, which are often black\-box and non\-differentiable\. Our proposed optimization scheme builds upon the idea of noise optimization with several key novelties\. First, whereas prior methods operate only on continuous latent noise, we extend the optimization framework to handle the categorical sampling process for discrete atom types\. Second, while existing works mainly consider differentiable reward functions, we investigate practical solutions for optimization under non\-differentiable, black\-box evaluators so that the framework can be applied to real\-world scenarios\. Finally, to the best of our knowledge, this work is the first to explore noise optimization in the context of drug design, and we demonstrate that this can be an effective mechanism for aligning diffusion\-based generation of NCEs with specific objectives \(e\.g\., higher absorption and lower toxicity\)\.

## Materials

### Datasets

We use the training and testing sets of the widely\-adopted CrossDocked2020\[Francoeur2020\]benchmark, referred to as๐–ข๐–ฃ\\mathop\{\\mathsf\{CD\}\}\\limits, for model training and testing, respectively\. In addition, we also use a new, manually annotated dataset,๐–ข๐–ฃ๐–ง\\mathop\{\\mathsf\{CDH\}\}\\limits, for testing\.

#### Training data

We follow the data preparation and splitting process of Luo*et al\.*\[luo2021sbdd\], refining 22\.5 million docked binding complexes to high\-quality docking poses\. For each protein in๐–ข๐–ฃ\\mathop\{\\mathsf\{CD\}\}\\limits, there are multiple ligand\-binding complexes:\(1\)determined experimentally and reported in the Protein Data Bank \(PDB\)\[berman2000protein\], and\(2\)obtained through simulation by redocking ligands into non\-cognate receptors\. We retain high\-quality poses, including experimentally determined and docked poses whose heavy\-atom RMSD < 1 ร… relative to the experimentally\-observed ligands pose from the PDB, and select diverse proteins with pairwise sequence identity below 30%\. After filtering, the dataset consists of 100,000 protein\-ligand pairs for training\. As๐–ข๐–ฃ\\mathop\{\\mathsf\{CD\}\}\\limitsdoes not provide explicit pocket regions, following Guan*et al\.*\[guan2023targetdiff\], we define the binding pocket as all residues and their atoms within a 10 ร… radius of its ligand\.

#### Testing data

##### ๐–ข๐–ฃ\\mathop\{\\mathsf\{CD\}\}\\limitstesting data

The๐–ข๐–ฃ\\mathop\{\\mathsf\{CD\}\}\\limitstest set is consistent with the data preparation and splitting process described in๐–ข๐–ฃ\\mathop\{\\mathsf\{CD\}\}\\limitsโ€™s training data\. We use all 100 protein\-ligand complexes from๐–ข๐–ฃ\\mathop\{\\mathsf\{CD\}\}\\limitsโ€™s test set for testing\. We refer to the known ligand provided in the test set as the reference ligand\. The latter is used only to identify the pocket region in the protein structure\. Once the pocket is identified, it is provided to๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsand๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsas input, while the reference ligand itself is excluded from the generation process\. The test set includes human proteins associated with oncological, neurodegenerative, infectious, cardiovascular, and metabolic diseases, as well as non\-human proteins such as bacterial, fungal, parasitic, viral, and plant proteins\.

##### ๐–ข๐–ฃ๐–ง\\mathop\{\\mathsf\{CDH\}\}\\limitstesting data

While๐–ข๐–ฃ\\mathop\{\\mathsf\{CD\}\}\\limitsโ€™s test set has been widely used in the research community, we discovered that it contains protein targets for non\-human species, such as plants and bacteria, which limits its utility for human therapeutics discovery\. To address this limitation, we manually curated a new benchmarking dataset based on๐–ข๐–ฃ\\mathop\{\\mathsf\{CD\}\}\\limits, referred to as๐–ข๐–ฃ๐–ง\\mathop\{\\mathsf\{CDH\}\}\\limits, which contains only human targets\.๐–ข๐–ฃ๐–ง\\mathop\{\\mathsf\{CDH\}\}\\limitsretains all 33 human targets from๐–ข๐–ฃ\\mathop\{\\mathsf\{CD\}\}\\limitsโ€™s test set, and includes an additional 39 human targets \(9 orthologs, 30 new targets\) from the PDB\. These additional PDB targets are carefully selected to cover high\-burden diseases such as cancers, neurological disorders, and autoimmune and inflammatory diseases, making๐–ข๐–ฃ๐–ง\\mathop\{\\mathsf\{CDH\}\}\\limitsmore relevant to human diseases\. Table[E](https://arxiv.org/html/2607.12349#A5)in Appendix[E](https://arxiv.org/html/2607.12349#A5)presents the๐–ข๐–ฃ๐–ง\\mathop\{\\mathsf\{CDH\}\}\\limitsinformation\.

For all the targets in๐–ข๐–ฃ๐–ง\\mathop\{\\mathsf\{CDH\}\}\\limits, we consider their relevant key ADMET properties, including:\(1\)carcinogenicity \(Carci\), the propensity of a molecule to cause cancer,\(2\)Ames mutagenicity \(Ames\), the property of a molecule likely to mutate DNA,\(3\)hERG inhibition \(hERG\), the potential cardiac side effects due to inhibition of the cardiac hERG \(KCNH2\) potassium channel,\(4\)human intestinal absorption \(HIA\), the ability of an oral drug to be absorbed through the intestinal lining, and\(5\)blood\-brain barrier permeability \(BBBP\), the ability to cross the blood\-brain barrier and enter the central nervous system\. Assessment of carcinogenicity, Ames mutagenicity, and hERG inhibition is essential to ensure drug safety by mitigating risks of tumorigenicity, genotoxicity, and cardiotoxicity\. Evaluation of human intestinal absorption and blood\-brain barrier permeability is critical to establish pharmacokinetic suitability and therapeutic accessibility of drug candidates\. These ADMET properties represent key determinants in drug discovery, guiding the identification and optimization of viable therapeutic candidates\. Unfortunately, current generative SBDD models do not proactively evaluate or optimize these properties, largely missing opportunities to increase the overall success rate of the drug discovery and development process\.๐–ข๐–ฃ๐–ง\\mathop\{\\mathsf\{CDH\}\}\\limitsis constructed to enable such ADMET evaluation\.

### Baselines

To evaluate the effectiveness of๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsand๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsin generating ligands that bind to target protein pockets, we compare them against the following state\-of\-the\-art baselines for SBDD:๐– ๐–ฑ\\mathop\{\\mathsf\{AR\}\}\\limits\[luo2021sbdd\],๐–ฏ๐—ˆ๐–ผ๐—„๐–พ๐—๐Ÿค๐–ฌ๐—ˆ๐—…\\mathop\{\\mathsf\{Pocket2Mol\}\}\\limits\[peng22pocket2mol\],๐–ฃ๐—‚๐–ฟ๐–ฟ๐–ฒ๐–ก๐–ฃ๐–ฃ\\mathop\{\\mathsf\{DiffSBDD\}\}\\limits\[schneuing2022structure\],๐–ณ๐–บ๐—‹๐—€๐–พ๐—๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{TargetDiff\}\}\\limits\[guan2023targetdiff\],๐–ฃ๐–พ๐–ผ๐—ˆ๐—†๐—‰๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{DecompDiff\}\}\\limits\[guan2023decompdiff\],๐–ฃ๐–พ๐–ผ๐—ˆ๐—†๐—‰๐–ฎ๐—‰๐—\\mathop\{\\mathsf\{DecompOpt\}\}\\limits\[zhoudecompopt\],๐–จ๐–ฏ๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{IPDiff\}\}\\limits\[huang2024protein\]and๐– ๐–ซ๐–จ๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{ALIDiff\}\}\\limits\[gu2024aligning\], where๐– ๐–ฑ\\mathop\{\\mathsf\{AR\}\}\\limitsand๐–ฏ๐—ˆ๐–ผ๐—„๐–พ๐—๐Ÿค๐–ฌ๐—ˆ๐—…\\mathop\{\\mathsf\{Pocket2Mol\}\}\\limitsare non\-diffusion methods, and the rest are all diffusion\-based methods\. These methods generate 3D binding ligands conditioned on the pockets of protein targets\. We choose these baselines because they are well\-established SBDD methods with strong performance on๐–ข๐–ฃ\\mathop\{\\mathsf\{CD\}\}\\limits\. We used their author\-provided implementations with released checkpoints and default inference settings for all baselines to conduct computational experiments on๐–ข๐–ฃ๐–ง\\mathop\{\\mathsf\{CDH\}\}\\limits\(except๐– ๐–ซ๐–จ๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{ALIDiff\}\}\\limits, which does not provide checkpoints\)\.

### Model Training and Evaluation

All the models, including๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limits,๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits, and the baselines, are trained over๐–ข๐–ฃ\\mathop\{\\mathsf\{CD\}\}\\limitstraining data, and tested on the๐–ข๐–ฃ\\mathop\{\\mathsf\{CD\}\}\\limitstest set and๐–ข๐–ฃ๐–ง\\mathop\{\\mathsf\{CDH\}\}\\limits\. In๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limits,๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limitsand๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{pcLG\}\}\\limits$\}\}\\limitshave different training objectives, thus,๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limitsis trained using only the protein pockets from๐–ข๐–ฃ\\mathop\{\\mathsf\{CD\}\}\\limitstraining set, and๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{pcLG\}\}\\limits$\}\}\\limitsis trained using the protein\-ligand complexes\. In๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits, we focus on the targets in๐–ข๐–ฃ๐–ง\\mathop\{\\mathsf\{CDH\}\}\\limitsand their ADMET property optimization, and its generation\-time optimization does not need model finetuning\. For each target in๐–ข๐–ฃ๐–ง\\mathop\{\\mathsf\{CDH\}\}\\limits,๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsoptimizes one particular ADMET property key to that target during the generation time\. The properties to optimize for different disease targets are presented in Table[E](https://arxiv.org/html/2607.12349#A5)in Appendix[E](https://arxiv.org/html/2607.12349#A5)\. Thus, the evaluation is organized with respect to each ADMET property across multiple relevant targets\.

### Evaluation Metrics

#### General ligand properties

To predict binding affinity, following the literature\[guan2023targetdiff,guan2023decompdiff\], we use Vina Scores \(Vina S\) calculated by AutoDock Vina\[Eberhardt2021\], which evaluates the quality of the binding poses of generated ligands against protein targets\. In addition, as suggested in the literature\[guan2023targetdiff,guan2023decompdiff\], we optimize the poses of the generated 3D ligands using a local energy minimization algorithm and a docking algorithm, as implemented in AutoDock Vina\[Eberhardt2021\]\. We then evaluate the predicted binding affinities of these optimized poses via two metrics: Vina Minimization \(Vina M\) based on local energy minimization, and Vina Dock \(Vina D\) based on docking\. Lower Vina scores indicate stronger binding affinities\. Based on Vina D, following the literature\[guan2023targetdiff,guan2023decompdiff\], we also measure the percentage of how many generated ligands across all the targets bind better than their reference ligand, referred to as High Affinity percentage \(HA%\)\. Higher HA% indicates a better capacity to generate ligands above the reference affinities\.

For drug\-likeness, we evaluate whether the generated ligands are drug\-like using the quantitative estimate of drug\-likeness \(QED\)\[Bickerton2012\]and synthesizable using a synthetic accessibility \(SA\) score\[Ertl2009\]\. We also calculate the diversity among generated ligands of each target, which is a per\-target measurement and defined as the average pairwise Tanimoto distances\[bajusz2015tanimoto\]derived using 2048\-bit fingerprints as implemented within RDKit\[rdkit\]\. Higher diversity indicates a better ability to explore broader chemical space\. To jointly consider predicted binding affinity, drug\-likeness, and synthetic feasibility, following previous work\[guan2023targetdiff,guan2023decompdiff\], we evaluate the success rate \(SR%\) calculated as the percentage of all generated ligands across all targets with Vina D<<\-8\.18, QED\>\>0\.25, and SA\>\>0\.59\. Higher SR% indicates a better ability to generate high\-quality ligands\. The Vina D score threshold \-8\.18, which corresponds to a binding affinity of less than 1ฮผ\\muM, is widely recognized in medicinal chemistry as an indicator of moderate biological activity\[zhoudecompopt\]\. The QED and SA score thresholds of 0\.25 and 0\.59 are set as the 10th percentile of approved drugs in DrugCentral\[ursu2016drugcentral\], enforcing reasonable lower bounds to maximize the consideration of potentially promising ligands\.

#### ADMET properties

We evaluate the generated ligands on five ADMET properties: Carci, Ames, hERG, HIA, and BBBP\. We useADMET\-AI\[swanson2024admet\]to estimate a probabilistic score for each property\. For Carci, Ames, and hERG, lower scores are preferred, indicating lower toxicity risks, while for HIA and BBBP, higher scores are more desired, indicating higher human intestinal absorption and blood\-brain barrier permeability, respectively\.

#### Ligand structures

We also evaluate the quality of generated ligands by analyzing their stability and 3D structural quality\. Specifically, to evaluate stability, we calculate both atomic level and molecular level stability, as described in literature\[hoogeboom22diff\]\. Atomic level stability is defined by the percentage of atoms that maintain correct valency, while molecular level stability measures the percentage of molecules in which all the atoms are stable, both for all the generated ligands and for all the targets\. We evaluate the 3D structures of generated ligands using the same metrics as in Peng*et al\.*\[peng2023moldiff\]\. We calculate the Jensen\-Shannon \(JS\) divergences for bond lengths, bond angles, and dihedral angles, which measure how far the distributions of these properties of generated ligands are compared with those of real molecules \(i\.e\., training molecules\) regarding the 3D structures\.

## Results

We assess the performance of the baselines and๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitson the 100 test pockets from๐–ข๐–ฃ\\mathop\{\\mathsf\{CD\}\}\\limitstest set and the 72 curated test pockets from๐–ข๐–ฃ๐–ง\\mathop\{\\mathsf\{CDH\}\}\\limits, using the evaluation metrics with 100 ligands generated per pocket by these methods\. We evaluate๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsusing the๐–ข๐–ฃ๐–ง\\mathop\{\\mathsf\{CDH\}\}\\limits, optimizing the 100 ligands per pocket with respect to the ADMET properties most relevant to its corresponding therapeutic indication\. Here, we present the results on๐–ข๐–ฃ๐–ง\\mathop\{\\mathsf\{CDH\}\}\\limits; the results on๐–ข๐–ฃ\\mathop\{\\mathsf\{CD\}\}\\limitsare discussed in AppendixLABEL:supp:cd\.

### Overall Comparison on๐–ข๐–ฃ๐–ง\\mathop\{\\mathsf\{CDH\}\}\\limits

Table 1:Comparison on๐–ข๐–ฃ๐–ง\\mathop\{\\mathsf\{CDH\}\}\\limitsMethodVina Sโ†“\\downarrowVina Mโ†“\\downarrowVina Dโ†“\\downarrowHA%โ†‘\\uparrowQEDโ†‘\\uparrowSAโ†‘\\uparrowDivโ†‘\\uparrowSR%โ†‘\\uparrowAvg\.Med\.Avg\.Med\.Avg\.Med\.Avg\.Med\.Avg\.Med\.Avg\.Med\.Avg\.Med\.Reference\-7\.45\-7\.35\-7\.89\-7\.63\-8\.01\-7\.71\-\-0\.460\.460\.730\.76\-\-29\.2๐– ๐–ฑ\\mathop\{\\mathsf\{AR\}\}\\limits\-6\.63\-6\.51\-6\.95\-6\.69\-7\.43\-7\.2137\.021\.60\.520\.530\.630\.630\.690\.6912\.2๐–ฏ๐—ˆ๐–ผ๐—„๐–พ๐—๐Ÿค๐–ฌ๐—ˆ๐—…\\mathop\{\\mathsf\{Pocket2Mol\}\}\\limits\-5\.28\-5\.02\-6\.44\-6\.16\-7\.23\-7\.0033\.115\.70\.610\.610\.790\.800\.790\.8129\.7๐–ฃ๐—‚๐–ฟ๐–ฟ๐–ฒ๐–ก๐–ฃ๐–ฃ\\mathop\{\\mathsf\{DiffSBDD\}\}\\limits\-3\.24\-5\.12\-5\.27\-5\.94\-7\.02\-7\.3034\.121\.90\.480\.490\.630\.610\.780\.7511\.3๐–ณ๐–บ๐—‹๐—€๐–พ๐—๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{TargetDiff\}\}\\limits\-6\.38\-6\.67\-7\.18\-7\.25\-8\.24\-8\.1947\.143\.90\.460\.470\.580\.570\.710\.7012\.3๐–ฃ๐–พ๐–ผ๐—ˆ๐—†๐—‰๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{DecompDiff\}\}\\limits\-6\.17\-6\.02\-6\.96\-6\.76\-7\.88\-7\.8349\.845\.50\.480\.480\.640\.630\.670\.6618\.3๐–ฃ๐–พ๐–ผ๐—ˆ๐—†๐—‰๐–ฎ๐—‰๐—\\mathop\{\\mathsf\{DecompOpt\}\}\\limits\-6\.34\-6\.22\-7\.04\-6\.88\-7\.98\-7\.8351\.752\.90\.480\.480\.650\.640\.670\.6720\.5๐–จ๐–ฏ๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{IPDiff\}\}\\limits\-8\.02\-8\.15\-8\.43\-8\.30\-9\.24\-8\.9162\.066\.30\.460\.460\.550\.540\.720\.7110\.0๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limits\-6\.80\-7\.10\-7\.69\-7\.78\-8\.85\-8\.7162\.274\.60\.460\.450\.590\.580\.600\.5822\.0
- โ€ขโ€‹โ€‹Columns represent: โ€œVina Sโ€: the predicted binding affinities between the initially generated poses of ligands and the protein pockets; โ€œVina Mโ€: the predicted binding affinities between the poses after local structure minimization and the protein pockets; โ€œVina Dโ€: the predicted binding affinities between the poses determined by AutoDock Vina and the protein pockets; โ€œHAโ€: the percentage of generated ligands with Vina D lower than those of reference ligands; โ€œQEDโ€: the quantitative estimate of drug\-likeness; โ€œSAโ€: the synthetic accessibility score; โ€œDivโ€: the diversity among generated ligands; โ€œSRโ€: the percentage of generated molecules with predicted binding affinities, QED and SA above certain thresholds\. The best values are inbold, and the second\-best values areunderlined\. Rows ingrayare excluded from the comparison\.โ†‘\\uparrow/โ†“\\downarrowindicate that higher / lower values are better\.

Table[1](https://arxiv.org/html/2607.12349#Sx4.T1)presents a comprehensive comparison between๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsand all the baselines in terms of generating drug\-like and diverse ligands that effectively bind to protein pockets of๐–ข๐–ฃ๐–ง\\mathop\{\\mathsf\{CDH\}\}\\limits\.๐–จ๐–ฏ๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{IPDiff\}\}\\limitsis reported for completeness but excluded from best/second\-best marking because ligands generated by this method exhibit simplified structures with excessive carbons and single bonds \(Details are in the Analysis of Atom and Bond Composition on๐–ข๐–ฃ๐–ง\\mathop\{\\mathsf\{CDH\}\}\\limitssection\)\. In terms of Vina S \(Avg/Med\),๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsoutperforms all the baseline methods, and also achieves the highest HA%, with a 2\.6% and 24\.9% improvement over the best baselines on Vina S and HA%, respectively\. This suggests that๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsbetter guides ligand atoms toward regions where they can engage in interactions with the pocket\. After energy minimization and re\-docking,๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsremains superior in terms of Vina M and Vina D scores, with 7\.1% and 7\.4% improvement over the best baselines, respectively\. This indicates the generated initial poses offer strong starting points for further refinement\.๐– ๐–ฑ\\mathop\{\\mathsf\{AR\}\}\\limitsis the second\-best model in terms of Vina S\.๐– ๐–ฑ\\mathop\{\\mathsf\{AR\}\}\\limitslearns pocket\-anchored atom type density maps, making it easier to place atoms and achieve high binding affinities\. In terms of Vina M and Vina D,๐–ณ๐–บ๐—‹๐—€๐–พ๐—๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{TargetDiff\}\}\\limitsachieves the second\-best performance\. Unlike๐– ๐–ฑ\\mathop\{\\mathsf\{AR\}\}\\limits, whose auto\-regressive sampling strategy suffers from high variance in pose quality,๐–ณ๐–บ๐—‹๐—€๐–พ๐—๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{TargetDiff\}\}\\limitsgenerates more consistent poses, improving Vina M and Vina D over๐– ๐–ฑ\\mathop\{\\mathsf\{AR\}\}\\limitsbut still underperforming๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limits\. Among the baseline methods,๐–ฃ๐–พ๐–ผ๐—ˆ๐—†๐—‰๐–ฎ๐—‰๐—\\mathop\{\\mathsf\{DecompOpt\}\}\\limitsspecifically performs iterative, docking\-guided arm optimization, aiming to achieve high binding affinities and thus high Vina scores\. Instead,๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsdoes not explicitly conduct such optimization โ€“ the fact that it still outperforms๐–ฃ๐–พ๐–ผ๐—ˆ๐—†๐—‰๐–ฎ๐—‰๐—\\mathop\{\\mathsf\{DecompOpt\}\}\\limitsin binding affinity scores indicates๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsencodes pocket information and pocket\-ligand interactions better than baselines\. Compared with diffusion models that do not perform any Vina\-based optimization \(๐–ฃ๐—‚๐–ฟ๐–ฟ๐–ฒ๐–ก๐–ฃ๐–ฃ\\mathop\{\\mathsf\{DiffSBDD\}\}\\limits,๐–ณ๐–บ๐—‹๐—€๐–พ๐—๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{TargetDiff\}\}\\limits,๐–ฃ๐–พ๐–ผ๐—ˆ๐—†๐—‰๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{DecompDiff\}\}\\limits\),๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsattains the best binding affinity scores while being more sample efficient\. The key difference is that๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsconditions generation on a pretrained pocket embedding, so pocket information is available*a priori*and does not need to be encoded every timestep during sampling, reducing overhead and improving pose quality\. In addition,๐–ฃ๐—‚๐–ฟ๐–ฟ๐–ฒ๐–ก๐–ฃ๐–ฃ\\mathop\{\\mathsf\{DiffSBDD\}\}\\limitsshows substantially lower Vina S than other methods\. A possible reason is coarse encoding of pocket structure, making๐–ฃ๐—‚๐–ฟ๐–ฟ๐–ฒ๐–ก๐–ฃ๐–ฃ\\mathop\{\\mathsf\{DiffSBDD\}\}\\limitsignore some important structural information and potential interaction sites\. This leads๐–ฃ๐—‚๐–ฟ๐–ฟ๐–ฒ๐–ก๐–ฃ๐–ฃ\\mathop\{\\mathsf\{DiffSBDD\}\}\\limitsto struggle to place ligands correctly within the pocket, causing steric clashes, suboptimal functional\-group orientations, and consequently higher Vina scores\. In contrast,๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsavoids this issue through its dedicated pocket encoder\.

The๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsmethod achieves QED and SA scores comparable with those of other diffusion baselines\. Notably, the two autoregressive models \(๐– ๐–ฑ\\mathop\{\\mathsf\{AR\}\}\\limitsand๐–ฏ๐—ˆ๐–ผ๐—„๐–พ๐—๐Ÿค๐–ฌ๐—ˆ๐—…\\mathop\{\\mathsf\{Pocket2Mol\}\}\\limits\) exhibit QED and SA scores significantly higher than those of all diffusion\-based models\. These two methods tend to generate smaller and simpler molecules via an atom\-by\-atom autoregression, with an average heavy\-atom count of 17 and 20, respectively, compared to 26 from other diffusion\-based baselines and 28 from the references\. While such molecules are favored by QED and SA measurement due to small size, simple structures and limited scaffolds\[bickerton2012quantifying,ertl2009estimation\],๐– ๐–ฑ\\mathop\{\\mathsf\{AR\}\}\\limitsand๐–ฏ๐—ˆ๐–ผ๐—„๐–พ๐—๐Ÿค๐–ฌ๐—ˆ๐—…\\mathop\{\\mathsf\{Pocket2Mol\}\}\\limitsdo not generally enjoy the same favor in terms of binding affinities\. Among the diffusion\-based methods, the decomposition baselines \(๐–ฃ๐–พ๐–ผ๐—ˆ๐—†๐—‰๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{DecompDiff\}\}\\limitsand๐–ฃ๐–พ๐–ผ๐—ˆ๐—†๐—‰๐–ฎ๐—‰๐—\\mathop\{\\mathsf\{DecompOpt\}\}\\limits\) achieve highest QED and SA scores\. Their decomposed, multiple priors for different atom roles bias the generation of molecules toward those with more drug\-like, synthesizable fragments, yielding higher QED and SA scores than those from other diffusion models of a single prior\.๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsemphasizes binding\-quality during generation, and therefore, shows moderate QED and SA with high binding affinities\.

We observe that๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsโ€‹โ€™s diversity slightly falls behind the baselines\. However,๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsโ€‹โ€™s diversity is still acceptable given its strengths in generating high\-affinity molecules, and it is comparable with those of๐–ฃ๐–พ๐–ผ๐—ˆ๐—†๐—‰๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{DecompDiff\}\}\\limitsand๐–ฃ๐–พ๐–ผ๐—ˆ๐—†๐—‰๐–ฎ๐—‰๐—\\mathop\{\\mathsf\{DecompOpt\}\}\\limits\. We hypothesize that the diversity is limited because the pocket is pre\-encoded as a fixed condition for diffusion\. A static pocket representation may sharpen the denoising landscape, limiting the chemical space that๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsmay visit\. In contrast, the baselines that jointly diffuse ligand and pocket allow the pocket representation to adapt to noisy states of ligands, and thus, achieve higher ligand diversity\. In terms of SR%,๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsattains the second\-best SR%, notably above all diffusion baselines, indicating that it balances design objectives of drugs effectively and yields realistic ligand candidates that are both effective binders and chemically viable\.๐–ฏ๐—ˆ๐–ผ๐—„๐–พ๐—๐Ÿค๐–ฌ๐—ˆ๐—…\\mathop\{\\mathsf\{Pocket2Mol\}\}\\limitsachieves the highest SR%, consistent with its strong QED and SA performance\. However,๐–ฏ๐—ˆ๐–ผ๐—„๐–พ๐—๐Ÿค๐–ฌ๐—ˆ๐—…\\mathop\{\\mathsf\{Pocket2Mol\}\}\\limitssignificantly underperforms๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsin terms of Vina scores\.

### Overall Comparison after ADMET Optimization

##### Overall Performance of๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits

Table 2:Comparison on๐–ข๐–ฃ๐–ง\\mathop\{\\mathsf\{CDH\}\}\\limitswith optimized property CarciMethodVina Sโ†“\\downarrowVina Mโ†“\\downarrowVina Dโ†“\\downarrowHA%โ†‘\\uparrowQEDโ†‘\\uparrowSAโ†‘\\uparrowDivโ†‘\\uparrowSR%โ†‘\\uparrowCarciโ†“\\downarrowAvg\.Med\.Avg\.Med\.Avg\.Med\.Avg\.Med\.Avg\.Med\.Avg\.Med\.Avg\.Med\.Avg\.Med\.Reference\-7\.53\-7\.45\-7\.93\-7\.68\-8\.00\-7\.71\-\-0\.450\.430\.730\.77\-\-25\.00\.190\.14๐– ๐–ฑ\\mathop\{\\mathsf\{AR\}\}\\limits\-6\.52\-6\.44\-6\.87\-6\.61\-7\.37\-7\.0936\.721\.00\.520\.530\.630\.630\.680\.6811\.20\.140\.07๐–ฏ๐—ˆ๐–ผ๐—„๐–พ๐—๐Ÿค๐–ฌ๐—ˆ๐—…\\mathop\{\\mathsf\{Pocket2Mol\}\}\\limits\-5\.30\-5\.04\-6\.46\-6\.15\-7\.25\-6\.9634\.514\.50\.610\.620\.790\.800\.790\.8029\.50\.180\.13๐–ฃ๐—‚๐–ฟ๐–ฟ๐–ฒ๐–ก๐–ฃ๐–ฃ\\mathop\{\\mathsf\{DiffSBDD\}\}\\limits\-3\.11\-5\.12\-5\.22\-5\.92\-6\.88\-7\.2534\.320\.30\.490\.500\.630\.620\.780\.7611\.60\.180\.12๐–ณ๐–บ๐—‹๐—€๐–พ๐—๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{TargetDiff\}\}\\limits\-6\.37\-6\.60\-7\.17\-7\.23\-8\.20\-8\.1547\.941\.10\.460\.470\.590\.580\.710\.6912\.30\.200\.15๐–ฃ๐–พ๐–ผ๐—ˆ๐—†๐—‰๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{DecompDiff\}\}\\limits\-6\.12\-5\.88\-6\.93\-6\.63\-7\.86\-7\.7448\.549\.50\.480\.490\.640\.620\.660\.6617\.10\.200\.15๐–ฃ๐–พ๐–ผ๐—ˆ๐—†๐—‰๐–ฎ๐—‰๐—\\mathop\{\\mathsf\{DecompOpt\}\}\\limits\-6\.30\-6\.14\-7\.02\-6\.79\-8\.00\-7\.7854\.156\.10\.480\.490\.640\.640\.670\.6619\.10\.200\.15๐–จ๐–ฏ๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{IPDiff\}\}\\limits\-7\.98\-8\.23\-8\.41\-8\.40\-9\.22\-8\.9162\.166\.30\.460\.470\.550\.540\.720\.7110\.20\.220\.18๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limits\-6\.66\-6\.80\-7\.46\-7\.45\-8\.48\-8\.4255\.962\.30\.490\.510\.620\.600\.650\.6223\.00\.220\.22๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-6\.55\-6\.80\-7\.35\-7\.38\-8\.37\-8\.3654\.256\.60\.520\.530\.600\.580\.630\.6118\.20\.060\.06- โ€ขโ€‹โ€‹Columns represent: โ€œVina Sโ€: the predicted binding affinities between the initially generated poses of ligands and the protein pockets; โ€œVina Mโ€: the predicted binding affinities between the poses after local structure minimization and the protein pockets; โ€œVina Dโ€: the predicted binding affinities between the poses determined by AutoDock Vina and the protein pockets; โ€œHAโ€: the percentage of generated ligands with Vina D lower than those of reference ligands; โ€œQEDโ€: the quantitative estimate of drug\-likeness; โ€œSAโ€: the synthetic accessibility score; โ€œDivโ€: the diversity among generated ligands; โ€œSRโ€: the percentage of generated molecules with binding affinities, QED, and SA above a certain threshold\. โ€œCarciโ€: the predicted carcinogenicity score for generated molecules; The best values are inbold, and the second\-best values areunderlined\. Rows ingrayare excluded from the comparison\.โ†‘\\uparrow/โ†“\\downarrowindicate that higher / lower values are better\.

The๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsmethod maintains strong performance on general molecular metrics compared with all evaluated baselines, ranking among the top methods in predicted binding affinity \(Vina S/M/D\), drug\-likeness \(QED\), and synthetic accessibility \(SA\)\. To examine๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsโ€™s performance in more detail, in Table[2](https://arxiv.org/html/2607.12349#Sx4.T2), we take the carcinogenicity optimization setting as an example, and compare the general metrics of๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsligands with those of๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsligands to ensure that the ADMET optimization procedure does not disrupt the desirable structures of initial๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsligands\. In general, the results show that ligands produced by๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsachieve Vina S/M/D, QED, and SA scores that are on par with, and in some cases exceed, those of๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limits\. As shown in Table[2](https://arxiv.org/html/2607.12349#Sx4.T2),๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsreduces the average Carci score from 0\.22 to 0\.06 \(by 73%\), which is the best among all methods, while the average Vina S/M/D scores change only slightly by less than 2\.5%\. Despite this marginal degradation, the resulting Vina scores of๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsremain stronger than all other baselines \(except๐–จ๐–ฏ๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{IPDiff\}\}\\limits\)\. The๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsand๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsmethods perform very similarly across QED and SA scores in Tables[2](https://arxiv.org/html/2607.12349#Sx4.T2), indicating that drug\-likeness and synthetic accessibility are well preserved after optimization\. This preservation can be attributed to the noise optimization design in๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits: rather than directly modifying the molecular structures of ligands generated by๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limits, the optimization is performed in the diffusion noise space, refining the generation process while keeping ligands within๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsโ€™s learned generative distribution\. We note that the diversity of๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsligands remains similar to that of๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsligands, suggesting that๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsenhances ADMET properties without collapsing the generated molecules into a narrow structural family\. We also summarize the results of optimizing other ADMET properties \(Ames, hERG, BBBP and HIA\) in TablesLABEL:tbl:overall\_results\_ames\-LABEL:tbl:overall\_results\_hiain AppendixLABEL:supp:opt\_results, and observe a similar pattern to carcinogenicity optimization:๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsconsistently improves the targeted ADMET score substantially while only modestly affecting binding affinity scores and other general molecular metrics\.

##### ADMET Property Optimization

Our results demonstrate that, leveraging noise optimization, the sampling process of๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsis guided towards a chemical subspace with more favorable ADMET properties, facilitating efficient discovery of safer and more effective drug candidates\. Table[3](https://arxiv.org/html/2607.12349#Sx4.T3)presents a comparison of generated molecules across different ADMET properties of interest, where each property score is optimized in a separate experiment\. The reported results of๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsare obtained by evaluating the best\-performing outcomes across multiple optimization iterations\. As the results show,๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsconsistently generates molecules with substantially more favorable ADMET properties compared to those from the baselines\. Specifically,๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsproduces ligands with the lowest predicted carcinogenicity and Ames mutagenicity, achieving roughly 50% to 60% better average Carci and Ames scores, compared to all evaluated baselines\. This indicates๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsgenerates drug candidates with much lower predicted risks of cancer induction and DNA damage across evaluated protein targets\. Similarly, Table[3](https://arxiv.org/html/2607.12349#Sx4.T3)demonstrates that๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsdecreases the hERG property score by 32%, relative to the next best baseline \(๐–ฏ๐—ˆ๐–ผ๐—„๐–พ๐—๐Ÿค๐–ฌ๐—ˆ๐—…\\mathop\{\\mathsf\{Pocket2Mol\}\}\\limits\), indicating considerable reduction in the predicted cardiotoxicity of the generated ligands\. Finally,๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsenhances absorption\-related properties, achieving notable improvements in both BBBP and HIA, compared to other baseline methods\. Taken together, these results highlight the effectiveness of๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsas an optimization method targeting desired ADMET objectives\.

Table 3:Comparison of ADMET property scores on๐–ข๐–ฃ๐–ง\\mathop\{\\mathsf\{CDH\}\}\\limitsMethodCarciโ†“\\downarrowAmesโ†“\\downarrowhERGโ†“\\downarrowBBBPโ†‘\\uparrowHIAโ†‘\\uparrowAvg\.Med\.Avg\.Med\.Avg\.Med\.Avg\.Med\.Avg\.Med\.Reference0\.190\.140\.330\.280\.440\.380\.660\.700\.730\.98๐– ๐–ฑ\\mathop\{\\mathsf\{AR\}\}\\limits0\.140\.070\.480\.470\.310\.210\.700\.770\.880\.99๐–ฏ๐—ˆ๐–ผ๐—„๐–พ๐—๐Ÿค๐–ฌ๐—ˆ๐—…\\mathop\{\\mathsf\{Pocket2Mol\}\}\\limits0\.180\.130\.420\.400\.220\.110\.760\.840\.941\.00๐–ฃ๐—‚๐–ฟ๐–ฟ๐–ฒ๐–ก๐–ฃ๐–ฃ\\mathop\{\\mathsf\{DiffSBDD\}\}\\limits0\.180\.120\.460\.430\.300\.210\.710\.770\.911\.00๐–ณ๐–บ๐—‹๐—€๐–พ๐—๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{TargetDiff\}\}\\limits0\.200\.150\.360\.290\.350\.260\.680\.720\.921\.00๐–ฃ๐–พ๐–ผ๐—ˆ๐—†๐—‰๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{DecompDiff\}\}\\limits0\.200\.150\.340\.270\.380\.280\.700\.750\.820\.99๐–ฃ๐–พ๐–ผ๐—ˆ๐—†๐—‰๐–ฎ๐—‰๐—\\mathop\{\\mathsf\{DecompOpt\}\}\\limits0\.200\.150\.350\.290\.370\.260\.690\.750\.821\.00๐–จ๐–ฏ๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{IPDiff\}\}\\limits0\.220\.180\.240\.160\.430\.460\.800\.860\.971\.00๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limits0\.220\.220\.370\.380\.430\.440\.640\.620\.930\.98๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits0\.060\.060\.100\.090\.150\.120\.890\.900\.991\.00- โ€ขColumns represent: โ€œCarciโ€: the carcinogenicity; โ€œAmesโ€: Ames mutagenicity; โ€œhERGโ€: hERG inhibition; โ€œBBBPโ€: blood\-brain barrier permeability; โ€œHIAโ€: human intestinal absorption\. The best values are inbold, and the second\-best values areunderlined\.โ†‘\\uparrow/โ†“\\downarrowindicate that higher / lower values are better\.

##### Change in ADMET Properties over Optimization Iterations

To gain deeper insight into how the optimized as well as other non\-optimized ADMET properties are affected during the optimization process, in Figure[2](https://arxiv.org/html/2607.12349#Sx4.F2), we visualize the changes in ADMET property scores of๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsgenerated ligands across optimization iterations\. Each property in the corresponding subplot is optimized independently, while the trajectories of non\-optimized properties are tracked in parallel\. For optimized properties, the scores consistently shift in the intended direction of improvement as the number of iterations increases โ€“ carcinogenicity, Ames, and hERG scores gradually decline, while HIA and BBBP scores steadily increase\. In contrast, the trajectories of non\-optimized properties remain largely stable, indicating that, in๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits, optimizing one property does not substantially compromise other properties\. Nevertheless, we observe mild tradeoffs across certain properties\. For instance, reducing hERG tends to slightly lower BBBP and HIA scores, and conversely, improving absorption\-related properties \(HIA, BBBP\) can increase hERG\. This coupling reflects common physicochemical and structural factors underlying these ADMET properties\. A moleculeโ€™s ability to permeate intestinal and blood\-brain barriers is often influenced by its lipophilicity, polarity, and molecular size\[weiss2024balanced\]\. However, these same factors also impact a moleculeโ€™s ability to bind the hERG potassium channel\. This highlights a key challenge in lead optimization: each structural modification has myriad effects that must be carefully monitored and controlled\. Disentangling such effects remains a difficult problem, which we leave for future exploration, possibly through multi\-objective optimization\.

![Refer to caption](https://arxiv.org/html/2607.12349v1/x2.png)

Fig\. 2:Changes in ADMET property scores over๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsoptimization iterations\. In each subfigure, the solid line indicates the property being optimized, while the dashed lines indicate the other properties not targeted for optimization\.๐š\\mathbf\{a\}\-๐ž\\mathbf\{e\}show the trends when optimizing Carci, Ames, hERG, HIA, and BBBP, respectively\.

### Comparison on Molecular Structures of Generated Ligands

Table 4:Comparison on molecular structures of generated ligands on๐–ข๐–ฃ๐–ง\\mathop\{\\mathsf\{CDH\}\}\\limitsGroupMetric๐– ๐–ฑ\\mathop\{\\mathsf\{AR\}\}\\limits๐–ฏ๐—ˆ๐–ผ๐—„๐–พ๐—๐Ÿค๐–ฌ๐—ˆ๐—…\\mathop\{\\mathsf\{Pocket2Mol\}\}\\limits๐–ฃ๐—‚๐–ฟ๐–ฟ๐–ฒ๐–ก๐–ฃ๐–ฃ\\mathop\{\\mathsf\{DiffSBDD\}\}\\limits๐–ณ๐–บ๐—‹๐—€๐–พ๐—๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{TargetDiff\}\}\\limits๐–ฃ๐–พ๐–ผ๐—ˆ๐—†๐—‰๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{DecompDiff\}\}\\limits๐–จ๐–ฏ๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{IPDiff\}\}\\limits๐–ฃ๐–พ๐–ผ๐—ˆ๐—†๐—‰๐–ฎ๐—‰๐—\\mathop\{\\mathsf\{DecompOpt\}\}\\limits๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limits๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsStabilityAtom stability \(โ†‘\\uparrow\)0\.9220\.8710\.8890\.9490\.9060\.9390\.9160\.9350\.937Molecule stability \(โ†‘\\uparrow\)0\.4490\.1710\.2200\.3890\.2350\.3560\.2670\.3300\.3493D structuresJS\. bond lengths \(โ†“\\downarrow\)0\.4630\.4370\.3290\.3360\.2980\.4810\.2850\.2720\.272JS\. bond angles \(โ†“\\downarrow\)0\.3720\.2440\.2980\.2340\.1640\.3200\.1590\.1870\.190JS\. dihedral angles \(โ†“\\downarrow\)0\.4080\.2330\.2900\.2580\.2020\.3250\.1990\.1460\.147

- Rows represent: โ€œatom stabilityโ€: the proportion of stable atoms that have the correct valency; โ€œmolecule stabilityโ€: the proportion of generated ligands with all atoms stable; โ€œJS\. bond lengths/bond angles/dihedral anglesโ€: the Jensenโ€“Shannon \(JS\) divergences of bond lengths, bond angles, and dihedral angles between generated ligands and training ligands\. The best values are inbold, and the second\-best values areunderlined\.โ†‘\\uparrow/โ†“\\downarrowindicate that higher / lower values are better\.

Table[4](https://arxiv.org/html/2607.12349#Sx4.T4)presents stability and 3D structure metrics for generated ligands, evaluating structural plausibility, such as proper atomic valences and realistic bond angles and dihedral angles geometry\. Notably,๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsand๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsachieve atom\-stability scores of 0\.935 and 0\.937, respectively, indicating strong adherence to correct valency at the atom level\. The๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsand๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsmethods also deliver competitive molecule stability, suggesting their ligands are chemically plausible\. Interestingly,๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsintroduces additional gains over๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitson both atom and molecule stability, suggesting the guidance from ADMET properties also biases generation towards more realistic structures with fewer valence violations and lower bond strain\.

Table[4](https://arxiv.org/html/2607.12349#Sx4.T4)also presents Jensen\-Shannon \(JS\) divergences\[lin2002divergence\]between the generated ligands and the training data ligands for bond lengths, bond angles, and dihedral angles\. The๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsand๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsmethods generate ligands with the lowest divergences from the training ligand distribution for bond lengths and dihedral angles, indicating both methods can generate molecules with realistic chemical structures\. Furthermore, the bond angle distributions for๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsand๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsare only slightly more divergent from the training distribution than the best baseline,๐–ฃ๐–พ๐–ผ๐—ˆ๐—†๐—‰๐–ฎ๐—‰๐—\\mathop\{\\mathsf\{DecompOpt\}\}\\limits\. Overall,๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsand๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsdemonstrate a strong ability to generate ligands with high\-quality chemical structures, mimicking the real\-world structural qualities of the training ligands\.

![Refer to caption](https://arxiv.org/html/2607.12349v1/x3.png)Fig\. 3:Comparison on filter pass rates of generated ligands to marketed drugs from ChemBL\[gaulton2012chembl\]\. Rule of Five, Rule of Ghose, Rule of Veber, and Rule of ZINC are common drug\-likeness rules, covering physicochemical properties, lipophilicity, hydrogen\-bonding capacity, molecular size, flexibility, and bioavailability\-related characteristics; BMS, PAINS, SureChEMBL, and NIBR are pharmaceutical structural alerts, covering undesirable, reactive, promiscuous, or assay\-interfering substructures; Complexity, Bredt, and Molecular Graph are cheminformatics validity tests, covering molecular complexity, strained structural motifs, and chemically implausible or unstable molecular graphs\.Figure[3](https://arxiv.org/html/2607.12349#Sx4.F3)presents the performance of generated ligands from๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsand๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitson several drug discovery filters, applying common drug\-likeness rules \(Rule of Five\[lipinski2004lead\], Rule of Ghose\[ghose1999knowledge\], Rule of Veber\[veber2002molecular\], Rule of ZINC\[irwin2005zinc,sterling2015zinc,irwin2020zinc20\]\), pharmaceutical structural alerts \(BMS\[pearce2006empirical\], PAINS\[baell2010new\], SureChEMBL\[papadatos2016surechembl\], NIBR\[schuffenhauer2020evolution\]\), and cheminformatics validity tests \(Complexity\[bertz1981first\], Bredt\[kobrich1973bredt\], Molecular Graph Validity\[polykovskiy2020molecular\]\)\. We select these filters as representative examples of commonly\-used filters in conventional drug discovery\[mignani2018present,kralj2023molecular\], removing undesirable chemical moieties and capturing a variety of perspectives on the properties of drug\-like molecules collected by chemists over decades of research\. Please note, these filters are not absolute determinants of success in drug discovery, but provide intuitive guidelines that reflect the real\-world experience of medicinal chemists\. Filter pass rates of marketed drugs are also reported, providing a fair external benchmark for general medicinal chemistry quality\. Note that these filter pass rates differ from the property metrics of Table[4](https://arxiv.org/html/2607.12349#Sx4.T4): Filter pass rates estimate rule\-based drug\-likeness and screen for known alert substructures, whereas Table[4](https://arxiv.org/html/2607.12349#Sx4.T4)verifies the chemical validity of molecules independent of drug\-likeness requirements\.

Ligands generated by๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsand๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitspass each filter at rates competitive with or higher than marketed drugs, indicating that they mirror the properties of approved drugs\. This suggests the generated ligands are largely compatible with conventional drug discovery wisdom, as reflected by these filters and rules\. Notably,๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsimproves the pass rates over๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitson pharmaceutical structural alerts \(BMS, PAINS, SureChEMBL, NIBR\), demonstrating its capability to reduce undesirable toxicophoric and reactive motifs through ADMET\-oriented optimization\. Overall, these filters highlight the promising nature of ligands generated from๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limits, while showing that๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitscan generally improve or maintain this performance, avoiding undesirable chemical moieties without harming the overall quality of generated ligands\.

### Analysis of Atom and Bond Composition on๐–ข๐–ฃ๐–ง\\mathop\{\\mathsf\{CDH\}\}\\limits

To evaluate the capability of the models in generating realistic molecular structures, it is essential to analyze the composition of atoms and bonds in generated ligands for chemical feasibility\. As shown in Figure[4](https://arxiv.org/html/2607.12349#Sx4.F4), in general,๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsand๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsproduce atom and bond compositions that are aligned with the reference distribution in training data of๐–ข๐–ฃ\\mathop\{\\mathsf\{CD\}\}\\limits\. Compared with the reference, all models show a slight overproduction of carbon\. This occurs because carbon dominates the training data \(67\.2%\), so the models trained on it tend to sample it more frequently\. Consequently, other atom types are slightly under sampled\. This pattern also appears in the distribution of bond types, where single bonds occur more often than in the training data \(52\.4%\)\. Among all bond types, aromatic rings require the most strict topological and geometric constraints\. Diffusion\-based models jointly denoise atom positions and types by focusing on local rather than global structures, which makes them prone to losing fine\-grained substructures, such as aromatic rings\. In contrast,๐–ฏ๐—ˆ๐–ผ๐—„๐–พ๐—๐Ÿค๐–ฌ๐—ˆ๐—…\\mathop\{\\mathsf\{Pocket2Mol\}\}\\limitspredicts atoms and bonds sequentially conditioned on previously sampled structure\. The bond prediction and the dependencies on previously generated structures allow it to better capture aromatic patterns\.

![Refer to caption](https://arxiv.org/html/2607.12349v1/x4.png)Fig\. 4:Comparison of atom \(a\-dfor C, N, O, and others \(F, S, Cl, P\), respectively\) and bond \(e\-hfor Single, Double, Trible, Aromatic, respectively\) type distributions across different models\. The dashed line displays the percentages in the training data as a realistic reference for drug molecules\.๐–จ๐–ฏ๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{IPDiff\}\}\\limitsachieves the best predicted binding affinity scores, as reported in TableLABEL:tbl:overall\_results\_crossdockand Table[1](https://arxiv.org/html/2607.12349#Sx4.T1)\. Its high affinity scores may stem from generating carbon\-heavy, single\-bond\-dominated structures with few heteroatoms and limited aromaticity, as Figure[4](https://arxiv.org/html/2607.12349#Sx4.F4)shows\. Excessive carbons in the ligand scaffold render the ligand broadly hydrophobic and prone to bind the pocket in a non\-specific way\. Although this may yield high Vina scores, it does not indicate a high intrinsic binding affinity\. The hydrophobic behavior also implies high lipophilicity, which often leads to poor ADMET properties and a higher risk of off\-target interactions\. At the same time, the scarcity of heteroatoms further lowers solubility and selectivity\. Therefore, although the ligands generated by๐–จ๐–ฏ๐–ฃ๐—‚๐–ฟ๐–ฟ\\mathop\{\\mathsf\{IPDiff\}\}\\limitsachieve high Vina scores, they are less attractive as drug candidates\.

### Case Studies

To demonstrate the practical utility of our methods, we use the programmed death\-ligand 1 \(PD\-L1\) and the colony\-stimulating factor\-1 receptor \(CSF1R\) targets for case studies\. PD\-L1 is a critical target for cancer immunotherapy due to its pivotal role in immune evasion by tumors\[mandal2025overcoming\]\. Current anti\-PD\-L1 therapies are being hampered by several clinical limitations, such as modest efficacy and resistance\[mandal2025overcoming\]\. Several high\-resolution crystal structures of PD\-L1\-ligand complexes are publicly available in the PDB database\[berman2000protein,Guzik2017\]\. CSF1R is a biologically important kinase involved in macrophage survival, proliferation, and differentiation\[wen2023csf1r,cannarile2017colony\]\. Current CSF1R inhibitors on the market, for example, pexidartinib and vimseltinib, have emerged as safe and efficacious options for the treatment of tenosynovial giant cell tumor \(TGCT\), a non\-malignant tumor of the joint, tendon sheath, or bursa driven by the overexpression of CSF1\. However, approved CSF1R\-targeting drugs still remain limited, motivating the development of additional candidate inhibitors\. CSF1R has a publicly available high\-quality co\-crystal structure\[kane2024identification\]\.

#### Experimental Setup

For both PD\-L1 and CSF1R, the case study pipeline consists of two sequential and complementary stages:\(1\)a structure\-based design stage to generate candidate ligands by๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits, and\(2\)a ligand\-based design stage to expand the๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated ligand set through analog search in existing corporate compounds deck and to prioritize promising compounds, thereby mimicking a hit expansion campaign as routinely done in early drug discovery\. These two stages together provide a rigorous workflow for candidate generation, screening, and selection prior to synthesis and subsequent biological testing\.

![Refer to caption](https://arxiv.org/html/2607.12349v1/x5.png)Fig\. 5:Docked poses of compound PL\-0 \(shown in magenta\), a known cognate ligand of PD\-L1, and compound PL\-1 \(shown in green\), a๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated novel ligand of PD\-L1, suggest that both compounds exhibit similar binding modes\.
![Refer to caption](https://arxiv.org/html/2607.12349v1/x6.png)Fig\. 6:Docked poses of compound CL\-0 \(shown in magenta\), a known cognate ligand of CSF1R, and compound CL\-12 \(shown in green\), a๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated novel ligand of CSF1R, suggest that both compounds exhibit similar binding modes\.

##### Structure\-Based Design via๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits

The co\-crystal structures of PD\-L1 and CSF1R bound to one of their respective cognate ligands are retrieved from the PDB\[berman2000protein\]under code 5N2F\[Guzik2017\]and 8W1L\[pdb8w1l,kane2024identification\], respectively\. The targets are prepared at pH 7\.4 using the protein wizard module as implemented in Maestro\[maestro2026\]\. Top\-ranked ligands generated by๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsin both cases were further filtered to remove undesirable chemical moieties as previously described by Kombo*et al\.*\[kombo2024predictions\]\. The remaining ligand structures are protonated at pH 7\.4 and energy\-minimized using Ligprep as implemented in Maestro\[maestro2026\]\. Flexible ligand docking was carried out using GLIDE in standard precision mode \(SP\), followed by extra precision mode \(XP\)\.

##### Ligand\-Based Design based on๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsGeneration

Following the structure\-based design stage, ligand\-based virtual screening of the Sanofi compound collection was carried out using FastROCS\[grant1996fast,rush2005shape,hawkins2007comparison\], a GPU\-accelerated tool for shape similarity search\. Ligands generated by๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsare used to build the shape query\. To further prioritize the identified virtual hits, machine learning \(ML\) models developed using structure\-activity relationship \(SAR\) datasets experimentally\-determined at Sanofi\[kane2024identification\]are used to predict potency against PD\-L1 and CSF1R, respectively\[kombo2024predictions\]\. Derived analogs are further scored and ranked using a weighted Pareto multi\-parameter optimization approach combining docking, ML\-predicted potency, and 3D similarity in shape and chemical features\.

#### Docking Results

Binding modes of reference compounds,๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated molecules, and their analogs are predicted by docking using GLIDE\[maestro2026\]\. Figure[6](https://arxiv.org/html/2607.12349#Sx4.F6)and[6](https://arxiv.org/html/2607.12349#Sx4.F6)show examples for PD\-L1 and CSF1R, respectively\. In the case of PD\-L1, the terminal phenyl group of the quintessential bi\-phenyl moiety in both the reference ligand and the ligand generated by๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsmakeฯ€\\pi\-ฯ€\\piinteractions with TYR\-56\. The other end of both ligands interacts with ASP\-122, one of the key residues driving molecular recognition at the PD\-L1 dimer interface\. In the case of CSF1R, both reference and๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated ligand make hydrogen\-bond interactions with CYS\-666 and ASP\-796, which constitute a hallmark of kinase inhibitors of this class\. These examples indicate that๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated ligands and their corresponding reference ligands have similar binding modes, further emphasizing the ability of๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsto learn from the binding site characteristics and accordingly generate high\-quality binding ligands\.

#### Biological Testing Results

![Refer to caption](https://arxiv.org/html/2607.12349v1/P2Diff_figures/two_cases/PDL1_structures/PL-0.png)PL\-0![Refer to caption](https://arxiv.org/html/2607.12349v1/P2Diff_figures/two_cases/PDL1_structures/PL-1.png)PL\-1![Refer to caption](https://arxiv.org/html/2607.12349v1/P2Diff_figures/two_cases/PDL1_structures/PL-2.png)PL\-2![Refer to caption](https://arxiv.org/html/2607.12349v1/P2Diff_figures/two_cases/PDL1_structures/PL-3.png)PL\-3![Refer to caption](https://arxiv.org/html/2607.12349v1/P2Diff_figures/two_cases/PDL1_structures/PL-4.png)PL\-4Fig\. 7:PD\-L1 ligand structures\. Each structure is labeled by its ID\.Table 5:SPR\-derived binding affinity and molecular properties of๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated PD\-L1 ligands and their analogs\.IDSourceKD\{\}\_\{\\text\{D\}\}\(ฮผ\\muM\)โ†“\\downarrowLogDQEDโ†‘\\uparrowLipEโ†‘\\uparrowPL\-1๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits3\.493\.010\.652\.45PL\-2๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits3\.753\.010\.652\.42PL\-3๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsanalog \(modified\)4\.002\.700\.642\.70PL\-4๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsanalog \(modified\)12\.300\.330\.504\.58PL\-0PDB: 5N2FND3\.040\.43ND
- โ€ขโ€‹โ€‹Columns represent: โ€œSourceโ€: the method used to obtain the molecule structure โ€“ โ€œmodifiedโ€ indicates the analog was manually modified from๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated molecules; โ€œKD\{\}\_\{\\text\{D\}\}โ€: SPR\-derived equilibrium dissociation constant inฮผ\\muM; โ€œLogDโ€: distribution coefficient, with a range between 1\-3 generally considered optimal in drug discovery; โ€œQEDโ€: quantitative estimate of drug\-likeness; โ€œLipEโ€: lipophilic efficiency;โ†‘\\uparrow/โ†“\\downarrowindicates that higher / lower values are better\.

![Refer to caption](https://arxiv.org/html/2607.12349v1/P2Diff_figures/two_cases/CSF1R_structures/CL-0.png)CL\-0![Refer to caption](https://arxiv.org/html/2607.12349v1/P2Diff_figures/two_cases/CSF1R_structures/CL-1.png)CL\-1![Refer to caption](https://arxiv.org/html/2607.12349v1/P2Diff_figures/two_cases/CSF1R_structures/CL-2.png)CL\-2![Refer to caption](https://arxiv.org/html/2607.12349v1/P2Diff_figures/two_cases/CSF1R_structures/CL-3.png)CL\-3![Refer to caption](https://arxiv.org/html/2607.12349v1/P2Diff_figures/two_cases/CSF1R_structures/CL-4.png)CL\-4![Refer to caption](https://arxiv.org/html/2607.12349v1/P2Diff_figures/two_cases/CSF1R_structures/CL-5.png)CL\-5![Refer to caption](https://arxiv.org/html/2607.12349v1/P2Diff_figures/two_cases/CSF1R_structures/CL-6.png)CL\-6![Refer to caption](https://arxiv.org/html/2607.12349v1/P2Diff_figures/two_cases/CSF1R_structures/CL-7.png)CL\-7![Refer to caption](https://arxiv.org/html/2607.12349v1/P2Diff_figures/two_cases/CSF1R_structures/CL-8.png)CL\-8![Refer to caption](https://arxiv.org/html/2607.12349v1/P2Diff_figures/two_cases/CSF1R_structures/CL-9.png)CL\-9![Refer to caption](https://arxiv.org/html/2607.12349v1/P2Diff_figures/two_cases/CSF1R_structures/CL-10.png)CL\-10![Refer to caption](https://arxiv.org/html/2607.12349v1/P2Diff_figures/two_cases/CSF1R_structures/CL-11.png)CL\-11![Refer to caption](https://arxiv.org/html/2607.12349v1/P2Diff_figures/two_cases/CSF1R_structures/CL-12.png)CL\-12Fig\. 8:CSF1R ligand structures\. Each structure is labeled by its ID\.Table 6:Kinase activity, molecular properties, and promiscuity profiling of๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated CSF1R molecules and their analogs\.IDSourceIC50\(nM\)โ†“\\downarrowLogDQEDโ†‘\\uparrowLipEโ†‘\\uparrow\#TargetsAssayed\#Target Hits\(ICโ‰ค5010ฮผ\{\}\_\{50\}\\leq 10~\\muM\)PromiscuityHit Rate \(%\)โ†“\\downarrowCL\-0PDB:8W1L201\.920\.365\.7836411\.1CL\-1๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsanalog \(VS\)2005\.060\.431\.649111\.1CL\-2๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsanalog \(VS\)3095\.750\.320\.7712632\.4CL\-3๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsanalog \(VS\)2,5005\.050\.460\.5512700CL\-4๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsanalog \(VS\)3,3632\.690\.482\.7917531\.7CL\-5๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsanalog \(VS\)3,7362\.140\.693\.298900CL\-6๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsanalog \(VS\)5,9391\.870\.423\.3511200CL\-7๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsanalog \(VS\)11,7002\.350\.632\.589500CL\-8๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsanalog \(VS\)21,2570\.930\.433\.742700CL\-9๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsanalog \(VS\)\>โ€‹30,000\\mathord\{\>\}30,0001\.240\.48ND9400CL\-10๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsInactive2\.820\.69NDโ€“โ€“โ€“CL\-11๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsInactive4\.820\.75NDโ€“โ€“โ€“CL\-12๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsND4\.660\.48NDโ€“โ€“โ€“
- โ€ขโ€‹โ€‹Columns represent: โ€œSourceโ€: the method used to obtain the molecule structure, where โ€œVSโ€ indicates the analog was identified via virtual screening with๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated molecules as shape query; โ€œIC50\(nM\)โ€: kinase activity against CSF1R measured in nM; โ€œLogDโ€: distribution coefficient, with a range between 1โ€“3 generally considered optimal in drug discovery; โ€œQEDโ€: quantitative estimate of drug\-likeness; โ€œLipEโ€: lipophilic efficiency; โ€œ\#Targets Assayedโ€: the number of distinct targets assayed in the selectivity or promiscuity panel; โ€œ\#Target Hitsโ€: the number of targets with ICโ‰ค5010ฮผ\{\}\_\{50\}\\leq 10~\\muM; โ€œPromiscuity Hit Rate \(%\)โ€: the percentage of assayed targets with ICโ‰ค5010ฮผ\{\}\_\{50\}\\leq 10~\\muM; โ€œInactiveโ€ indicates no significant inhibition was observed; โ€œNDโ€ indicates not determined;โ†‘\\uparrow/โ†“\\downarrowindicates that higher / lower values are better\.

Four compounds generated by๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\(PL\-1, PL\-2, CL\-10, CL\-11\), nine analogs derived from ligand\-based virtual screening of the Sanofi Corporate compound collection \(CL\-1, CL\-2, CL\-3, CL\-4, CL\-5, CL\-6, CL\-7, CL\-8, CL\-9\), and two manually modified analogs \(PL\-3, PL\-4\) have been synthesized \(AppendixLABEL:supp:synthesispresents the synthesis process\) and tested in biological assays \(AppendixLABEL:supp:assaypresents the biological testing protocols\)\. Lipophilic ligand efficiency \(LipE\) is calculated by subtracting calculated LogD at pH 7\.4 from pIC50and pKD\{\}\_\{\\text\{D\}\}in the case of CSF1R and PD\-L1 ligands, respectively\. LogD and QED were calculated using Pipeline Pilot\[pilot2020version\]\. The biological activity data for PD\-L1 and CSF1R are summarized in Tables[5](https://arxiv.org/html/2607.12349#Sx4.T5)and[6](https://arxiv.org/html/2607.12349#Sx4.T6), respectively; compound structures are presented in Figures[7](https://arxiv.org/html/2607.12349#Sx4.F7)and[8](https://arxiv.org/html/2607.12349#Sx4.F8), respectively\. AppendixLABEL:supp:ic50presents the dose\-response curves for CSF1R\.

For PD\-L1, as shown in Table[5](https://arxiv.org/html/2607.12349#Sx4.T5), the๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated ligands \(PL\-1, PL\-2\) and one of their modified analogs \(PL\-3\) exhibit low\-micromolar binding affinity, with values ranging from 3\.49 to 4\.00ฮผ\\muM and LipE values ranging from 2\.42 to 2\.70\. The other modified analog, PL\-4, exhibits binding at 12\.30ฮผ\\muM with a LipE of 4\.58\. It is encouraging to note that published mean LipE value for oral drugs is 4\.43\[hopkins2014role\], and LipE values increase from hit finding to drug candidate discovery\[kombo2025logic\]โ€“ Some๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated ligands already achieve ligand efficiency profiles comparable to established orally available drugs\. All four novel ligands exhibit QED values greater than the reference ligand \(PL\-0\)\.

As shown in Table[6](https://arxiv.org/html/2607.12349#Sx4.T6), results indicate that using๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated ligands as a template for building a 3D search query to find CSF1R ligands yielded two compounds, CL\-1 and CL\-2, with nanomolar kinase activity against CSF1R, with IC50values of 200 nM and 309 nM, respectively, and four other compounds, CL\-3, CL\-4, CL\-5, and CL\-6, with low micromolar kinase activity, with IC50values ranging from 2\.500 to 5\.939ฮผ\\muM\. Moreover, these compounds show low kinase promiscuity across the tested panels\. Specifically, under the 10ฮผ\\muM threshold, CL\-1 and CL\-2 have promiscuity hit rates of only 1\.1% and 2\.4%, respectively, which are substantially lower than the 11\.1% observed for the reference ligand \(CL\-0\)\. Among the other four compounds, all except CL\-4 show no detectable off\-target kinase hits under the same threshold, while CL\-4 shows only a low promiscuity hit rate of 1\.7%\. In addition, some of the virtual screening hits are known Factor Xa inhibitors, as shown by an expired patent\[ewing2001substituted\], which can inspire leveraging the observed polypharmacology to initiate compound repositioning studies\. Taken together, the results obtained in both target cases suggest that๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated ligands and their subsequent analogs have developability characteristics like those found in hit finding and hit expansion campaigns as carried out in drug discovery projects in pharmaceutical settings\. Furthermore, we have predicted that๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated ligands and their corresponding reference ligands have similar binding modes, further emphasizing the ability of๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitswhen it comes to efficiently deciphering the binding site characteristics\.

## Discussions and Conclusions

In this section, we discuss the limitations of๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsand๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsand sketch promising directions for future work\.

### Specializing Model for High\-selectivity Ligands

An important consideration in SBDD is ligand selectivity\. While current diffusion\-based models focus primarily on optimizing binding affinity to a given target, they often ignore off\-target interactions, which can lead to undesirable side effects\. A ligand that binds to both the intended target and unintended proteins may exhibit poor selectivity, reducing its therapeutic potential and increasing the risk of adverse effects\. To address this, future work could explore strategies for specializing generative models to generate highly selective ligands\. One approach is negative conditioning, training the model to avoid off\-targets while retaining on\-target affinity\. Alternatively, optimization techniques over on\- and off\-target scores can balance affinity and selectivity during generation\. These computational strategies could potentially improve the clinical relevance of structure\-based generated ligands and advance the applicability of such AI\-generative models in drug discovery\.

### Generalizing Model to Multi\-objective Optimization

While our current generation\-time optimization targets a single ADMET property at a time, real discovery tasks generally require simultaneous consideration of multiple ADMET endpoints, since strong affinity alone often fails to translate when even one critical ADMET criterion is not met\. A natural extension is a multi\-objective formulation that integrates multiple property signals simultaneously into the optimization module\. Such a method may construct composite update directions that coordinate improvements across all target properties or dynamically adapt the relative emphasis of each property according to which constraints are currently violated\. This would move the current single\-objective setting toward Pareto\-style multi\-objective optimization and enable fine\-grained control over affinity and ADMET trade\-offs while preserving the training\-free and plug\-and\-play nature of the optimization module\. For instance, AI\-MedCraft demonstrates how an adaptive Pareto\-guided strategy can coordinate multiple competing objectives\[barakat2026ai\]\. In practice, such flexibility could support different stages of discovery\. Tighter ADMET constraints could reduce downstream risk when optimizing near\-deployable candidates, whereas looser constraints could encourage broader exploration in early ideation, followed by gradual annealing toward stricter feasibility as optimization progresses\. Another promising direction is to make the optimization uncertainty\-aware\. Property predictors are often noisy, susceptible to distribution shift, and potentially vulnerable to exploitation under direct optimization\. To address this, the objective could incorporate risk\-sensitive criteria, such as lower\-confidence\-bound optimization or explicit penalties on predictive uncertainty, thereby steering updates toward candidates that are not only high\-scoring but also more reliable\. Together, these extensions would better align generative optimization with pharmacological and drug\-discovery practice, potentially accelerating ligand optimization toward clinic\-viable candidates\.

### Incorporating Chemical Synthesis Pathway into Diffusion Models

A major practical concern in drug discovery is the synthetic accessibility of proposed ligands\. While conDitar and conDitar\-dev generate chemically valid molecules and optimize the ADMET properties, they do not explicitly account for whether the binding ligand can be realized through feasible synthetic pathways\. As a result, some generated ligands, despite exhibiting good*in silico*profile, may be difficult to synthesize\. To address this limitation, future work could explore incorporating synthesis pathway information into diffusion models\. One possible direction is to introduce synthesis\-aware priors or guidance signals, such as retrosynthetic accessibility scores or reaction constraints, to bias generation toward synthetically accessible regions of chemical space\[shen2025compositional,cremer2024pilot\]\. Such a space could include commercially available chemical building blocks commonly used in chemical reactions carried out by medicinal chemists\. Alternatively, synthetic pathways could be modeled explicitly during diffusion training, enabling the diffusion process to simultaneously simulate molecular generation and the synthesis trajectory\. Such synthesis\-aware generative frameworks could enhance the alignment between generative model outputs and practical requirements in real\-world drug discovery workflows\.

### Informing Diffusion Models with Physical and Chemical Knowledge

Diffusion\-based methods may occasionally generate molecules with atypical ring systems, achieving high binding affinity scores while raising concerns regarding chemical realism\. To address this, one possible direction is to incorporate chemical domain knowledge and physical constraints into the generative process, thereby constructing a chemistry\- and physics\-informed diffusion framework\. For example, a framework like๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitscould be coupled with force field\-based energy minimization or other physics\-based refinement procedures to improve the structural realism of the generated ligands\. One such force field\-based method is OPLS\-AA\[jorgensen1988opls,jorgensen1996development,sambasivarao2009development,doherty2017revisiting\], which is widely used in molecular modeling and simulations of interactions between organic molecules and proteins\. Furthermore, in future work, the generated ligands and pocket residue side chains should be properly protonated at physiological pH \(7\.4\) prior to energy minimization and docking to ensure physically meaningful interaction modeling\. In addition to these improvements, we envision future work incorporating a robust protocol aimed at concomitantly filtering undesirable chemical moieties as ligands are being generated to steer the design process towards a more druggable chemical space\.

### Conclusions

We present๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits, a conditional diffusion\-based framework for SBDD capable of generating ligands with strong binding affinities and ADMET properties for drug discovery and development\. On๐–ข๐–ฃ๐–ง\\mathop\{\\mathsf\{CDH\}\}\\limits,๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsachieves better performance than the state\-of\-the\-art SBDD baselines, achieving an average Vina D score of \-8\.85 kcal/mol\. On five ADMET properties,๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsachieves improvements of up to 73% over๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitswhile retaining similar predicted binding affinity\. Furthermore, the molecules generated by๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsand๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsexhibit strong stability scores, realistic chemical structures, and high drug discovery filter pass rates\. Together, these results establish๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsas a strong framework for generating developable drug molecules\. Moreover,๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-designed molecules for CSF1R and PD\-L1, together with their derivatives and analogs, have been experimentally synthesized and biologically tested\. Several compounds show promising binding affinities in the nanomolar and low\-micromolar range, validating๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsbeyond the typical standard for*in silico*SBDD frameworks, which typically lack experimental validation\[hu2025target\]\. This case study demonstrates how๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitscan be used immediately by computational and medicinal chemists, highlighting its potential to transform drug discovery\. We envision๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsas a foundation for more advanced SBDD frameworks that incorporate additional developability objectives, including expanded ADMET properties, synthetic accessibility, structural constraints, and ligand selectivity\.

## Methods

We introduce๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits, a new framework for generating 3D ligands tailored to specific protein pockets, while also optimizing additional ADMET properties of the generated ligands\. As illustrated in Figure[1](https://arxiv.org/html/2607.12349#Sx1.F1),๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitscomprises three modules:\(1\)๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\\mathop\{\\mathsf\{msPRL\}\}\\limits, a pretrained pocket embedding module,\(2\)๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limits, a diffusion\-based, pocket\-conditioned ligand generation module, and\(3\)๐—‰๐–บ๐–ฎ๐–ฏ๐–ณ\\mathop\{\\mathsf\{paOPT\}\}\\limits, an optional optimization module applied at generation time of๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limits\.๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\\mathop\{\\mathsf\{msPRL\}\}\\limitslearns to represent 3D structures and physicochemical properties of protein binding pockets through an equivariant encoder\. For a given pocket, referred to as the condition pocket, its binding pocket representation produced by๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\\mathop\{\\mathsf\{msPRL\}\}\\limitswill be used in diffusion to generate new ligands\.๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsgenerates new, realistic ligands in 3D with high binding affinities to the condition pocket by learning from the pocket\-ligand interactions from existing complexes via a diffusion process\[ho2020ddpm\]\. To further improve other desired pharmacological properties of the generated ligands \(e\.g\., ADMET properties\[hodgson2001admet\]\),๐—‰๐–บ๐–ฎ๐–ฏ๐–ณ\\mathop\{\\mathsf\{paOPT\}\}\\limitsis built upon๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsby performing gradient\-based, zero*th*\-order optimization during its generation time without retraining๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limits, guiding the diffusion trajectories toward the desired properties\. Table[7](https://arxiv.org/html/2607.12349#Sx6.T7)presents the key notations used in the methods\. Particularly, subscriptsiandjindex atoms or residues\. Superscripts indicate the data domain:gfor ligand atoms,pfor pocket atoms,rfor pocket residues,cfor interactions between ligand atoms and pocket atoms, andafor interactions between ligand atoms and pocket residues\.

Table 7:Notationsnotationsmeanings๐’Ÿ\\mathop\{\\mathcal\{D\}\}\\limitsa ligand๐’ซ\\mathop\{\\mathcal\{P\}\}\\limitsa pocket๐–บi\\mathsf\{a\}\_\{i\}theii\-th atom in๐’Ÿ\\mathop\{\\mathcal\{D\}\}\\limits/๐’ซ\\mathop\{\\mathcal\{P\}\}\\limits๐—‹ip\\mathsf\{r\}^\{p\}\_\{i\}theii\-th residue in๐’ซ\\mathop\{\\mathcal\{P\}\}\\limits๐ฑi\\mathbf\{x\}\_\{i\}the position of๐–บi\\mathsf\{a\}\_\{i\}/๐—‹ip\\mathsf\{r\}^\{p\}\_\{i\}in 3D space๐ฏi\\mathbf\{v\}\_\{i\}the atom/residue type of๐–บi\\mathsf\{a\}\_\{i\}/๐—‹ip\\mathsf\{r\}^\{p\}\_\{i\}๐ฌi\\mathbf\{s\}\_\{i\}scalar embedding for atom๐–บi\\mathsf\{a\}\_\{i\}โ„‹i\\mathcal\{H\}\_\{i\}vector embedding for atom๐–บi\\mathsf\{a\}\_\{i\}dโ€‹\(๐–บi,๐–บj\)d\(\\mathsf\{a\}\_\{i\},\\mathsf\{a\}\_\{j\}\)the distance between๐–บi\\mathsf\{a\}\_\{i\}and๐–บj\\mathsf\{a\}\_\{j\}๐›iโ€‹j\\mbox\{$\\mathop\{\\mathbf\{b\}\}\\limits$\}\_\{ij\}the type of bond between๐–บig\\mathsf\{a\}^\{g\}\_\{i\}and๐–บjg\\mathsf\{a\}^\{g\}\_\{j\}gligand atomsppocket atomsrpocket residuescinteractions between ligand atoms and pocket atomsainteractions between ligand atoms and pocket residues### Preliminaries

Given the condition protein pocket๐’ซ\\mathop\{\\mathcal\{P\}\}\\limits,๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsโ€™s objective is to generate ligands exhibiting strong binding affinities to๐’ซ\\mathop\{\\mathcal\{P\}\}\\limitswith realistic 3D structures\.๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsextends this objective by optimizing additional physicochemical properties\. In this manuscript, we represent a ligand๐’Ÿ\\mathop\{\\mathcal\{D\}\}\\limitsas a set of its atoms, defined as

๐’Ÿ=\{๐–บ1g,๐–บ2g,โ€ฆ,๐–บNgโˆฃ๐–บig=\(๐ฑig,๐ฏig\)\},\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}=\\\{\\mathsf\{a\}^\{g\}\_\{1\},\\mathsf\{a\}^\{g\}\_\{2\},\\dots,\\mathsf\{a\}^\{g\}\_\{N\}\\mid\\mathsf\{a\}^\{g\}\_\{i\}=\(\\mathbf\{x\}^\{g\}\_\{i\},\\mathbf\{v\}^\{g\}\_\{i\}\)\\\},whereNNdenotes the number of atoms in๐’Ÿ\\mathop\{\\mathcal\{D\}\}\\limits\. Each atom๐–บig\\mathsf\{a\}^\{g\}\_\{i\}is characterized by its 3D coordinates๐ฑigโˆˆโ„1ร—3\\mathbf\{x\}^\{g\}\_\{i\}\\in\\mathbb\{R\}^\{1\\times 3\}and a one\-hot feature vector๐ฏigโˆˆโ„1ร—da\\mathbf\{v\}^\{g\}\_\{i\}\\in\\mathbb\{R\}^\{1\\times\{d\_\{a\}\}\}\. This feature vector encodes both the atom type and its aromaticity\. Here,da=13d\_\{a\}=13denotes the total number of distinct atom types considered \(e\.g\., carbon, oxygen, nitrogen; hydrogen is excluded\), including aromatic and non\-aromatic variants\. Similarly, we represent a protein pocket๐’ซ\\mathop\{\\mathcal\{P\}\}\\limitsusing its set of protein atoms within a predefined radius of its reference ligand \(as introduced in the section Datasets\) denoted asAp\\mathop\{\{A\}\_\{p\}\}\\limits, and the residues involved in the protein pocket, denoted asRp\\mathop\{\{R\}\_\{p\}\}\\limits, that is,

๐’ซ=\(Ap,Rp\),\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}=\(\\mbox\{$\\mathop\{\{A\}\_\{p\}\}\\limits$\},\\mbox\{$\\mathop\{\{R\}\_\{p\}\}\\limits$\}\),where

Ap=\{๐–บ1p,๐–บ2p,โ‹ฏ,๐–บNap\|๐–บip=\(๐ฑip,๐ฏip\)\},\\displaystyle=\\\{\\mathsf\{a\}^\{p\}\_\{1\},\\mathsf\{a\}^\{p\}\_\{2\},\\cdots,\\mathsf\{a\}^\{p\}\_\{N\_\{a\}\}\|\\mathsf\{a\}^\{p\}\_\{i\}=\(\\mathbf\{x\}^\{p\}\_\{i\},\\mathbf\{v\}^\{p\}\_\{i\}\)\\\},\(1\)Rp=\{๐—‹1p,๐—‹2p,โ‹ฏ,๐—‹Nrp\|๐—‹ip=\(๐ฑir,๐ฏir\)\}\.\\displaystyle=\\\{\\mathsf\{r\}^\{p\}\_\{1\},\\mathsf\{r\}^\{p\}\_\{2\},\\cdots,\\mathsf\{r\}^\{p\}\_\{N\_\{r\}\}\|\\mathsf\{r\}^\{p\}\_\{i\}=\(\\mathbf\{x\}^\{r\}\_\{i\},\\mathbf\{v\}^\{r\}\_\{i\}\)\\\}\.Each pocket atom๐–บipโˆˆAp\\mathsf\{a\}^\{p\}\_\{i\}\\in A\_\{p\}is characterized by its 3D coordinates๐ฑipโˆˆโ„1ร—3\\mathbf\{x\}^\{p\}\_\{i\}\\in\\mathbb\{R\}^\{1\\times 3\}and a one\-hot feature vector๐ฏipโˆˆโ„1ร—da\\mathbf\{v\}^\{p\}\_\{i\}\\in\\mathbb\{R\}^\{1\\times\{d\_\{a\}\}\}, representing its atom type withda=4d\_\{a\}=4\(carbon, nitrogen, oxygen, or sulfur\)\. Similarly, each pocket residue๐—‹ipโˆˆRp\\mathsf\{r\}^\{p\}\_\{i\}\\in R\_\{p\}is represented using the 3D coordinates of its alpha carbon atom๐ฑirโˆˆโ„1ร—3\\mathbf\{x\}^\{r\}\_\{i\}\\in\\mathbb\{R\}^\{1\\times 3\}, and a one\-hot feature vector๐ฏirโˆˆโ„1ร—20\\mathbf\{v\}^\{r\}\_\{i\}\\in\\mathbb\{R\}^\{1\\times\{20\}\}, representing the 20 amino acid types\. Note that each residue can be decomposed into its component atoms, represented as

๐—‹ip=\{๐–บi,1p,โ€ฆ,๐–บi,\|๐—‹ip\|p\},\\mathsf\{r\}^\{p\}\_\{i\}=\\\{\\mathsf\{a\}^\{p\}\_\{i,1\},\\dots,\\mathsf\{a\}^\{p\}\_\{i,\|\\mathsf\{r\}^\{p\}\_\{i\}\|\}\\\},where๐–บi,jp\\mathsf\{a\}^\{p\}\_\{i,j\}is thejj\-th atom in๐—‹ip\\mathsf\{r\}^\{p\}\_\{i\}\(๐–บi,jpโˆˆAp\\mathsf\{a\}^\{p\}\_\{i,j\}\\in A\_\{p\}\), and\|๐—‹ip\|\|\\mathsf\{r\}^\{p\}\_\{i\}\|is the number of atoms in๐—‹ip\\mathsf\{r\}^\{p\}\_\{i\}\. For each๐–บjpโˆˆAp\\mathsf\{a\}^\{p\}\_\{j\}\\in A\_\{p\}, the residue that it is in is denoted as๐—‹pโ€‹\(๐–บjp\)\\mathsf\{r\}^\{p\}\(\\mathsf\{a\}^\{p\}\_\{j\}\), that is,๐—‹pโ€‹\(๐–บi,jp\)=๐—‹ip\\mathsf\{r\}^\{p\}\(\\mathsf\{a\}^\{p\}\_\{i,j\}\)=\\mathsf\{r\}^\{p\}\_\{i\}\.

### Pocket Representation Pretraining \(๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\\mathop\{\\mathsf\{msPRL\}\}\\limits\)

We pretrain a pocket embedding module, denoted as๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\\mathop\{\\mathsf\{msPRL\}\}\\limits, to encode the 3D structures of target protein pockets into pocket embeddings\.๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\\mathop\{\\mathsf\{msPRL\}\}\\limitsemploys an encoder,๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits, which maps each pocket atom๐–บip\\mathsf\{a\}^\{p\}\_\{i\}into two latent representations: a scalar embedding๐ฌipโˆˆโ„d\\mathbf\{s\}^\{p\}\_\{i\}\\in\\mathbb\{R\}^\{d\}and a vector embeddingโ„‹ipโˆˆโ„dร—3\\mathcal\{H\}^\{p\}\_\{i\}\\in\\mathbb\{R\}^\{d\\times 3\}, whereddis a hyperparameter of๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\\mathop\{\\mathsf\{msPRL\}\}\\limits\. Together, these latent representations capture the essential structural information of the pocket\. Simultaneously,๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\\mathop\{\\mathsf\{msPRL\}\}\\limitsemploys a decoder,๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–ฝ๐–พ๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{dec\}\}\\limits, to decode pocket embeddings to atom positions and types\. FollowingUni\-Molโ€™s\[zhou2023uni\]approach,๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\\mathop\{\\mathsf\{msPRL\}\}\\limitsis optimized by reconstructing noise\-free pockets from corrupted ones using pocket embeddings, where the pockets are corrupted by adding uniform noise into the positions of randomly selected atoms with their atom types masked off\. Existing diffusion\-based methods\[guan2023targetdiff,guan2023decompdiff,huang2024protein\]typically jointly learn pocket representations and ligand representations together from pocket\-ligand interactions\. This approach can limit the quality of pocket representation learning since the comparatively larger size of pockets makes them more prone to being underrepresented or traded off relative to ligands\. In๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limits, we isolate the pocket representation learning within the pretraining of the๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\\mathop\{\\mathsf\{msPRL\}\}\\limitsmodule to ensure expressive pocket embeddings\.

#### Pocket Encoder \(๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits\)

๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limitslearns a scalar embedding๐ฌip\\mathbf\{s\}^\{p\}\_\{i\}and a vector embeddingโ„‹ip\\mathcal\{H\}^\{p\}\_\{i\}for each atom๐–บip\\mathsf\{a\}^\{p\}\_\{i\}in the pocket๐’ซ\\mathop\{\\mathcal\{P\}\}\\limitsto encode pocket structure and physicochemical properties\. These two embeddings integrate the information of each atom๐–บipโˆˆAp\\mathsf\{a\}^\{p\}\_\{i\}\\in A\_\{p\}and the residue๐—‹pโ€‹\(๐–บip\)โˆˆRp\\mathsf\{r\}^\{p\}\(\\mathsf\{a\}^\{p\}\_\{i\}\)\\in\\mbox\{$\\mathop\{\{R\}\_\{p\}\}\\limits$\}to which๐–บip\\mathsf\{a\}^\{p\}\_\{i\}belongs to capture multi\-granular information from the pockets\.

##### Pocket Atom Embeddings

๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limitsuses a multi\-layer graph attention neural network\[velickovic2018graph\]augmented with a geometric vector perceptron \(GVP\)\[jing2021learning\]to learn embeddings of pocket atoms\. At each layerllof the graph neural network,๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limitslearns the invariant features/embeddings๐ฌi,lp\\mathbf\{s\}^\{p\}\_\{i,l\}and the equivariant features/embeddingsโ„‹i,lp\\mathcal\{H\}^\{p\}\_\{i,l\}\(๐ฌi,0p=๐ฏip\\mathbf\{s\}^\{p\}\_\{i,0\}=\\mathbf\{v\}^\{p\}\_\{i\},โ„‹i,0p=๐ฑip\\mathcal\{H\}^\{p\}\_\{i,0\}=\\mathbf\{x\}^\{p\}\_\{i\}\)\. Invariant features capture scalar properties such as atom types and interatomic distances, which remain unchanged under rotations and translations; equivariant features encode vector information such as atom positions and spatial directions that are difference vectors between atom positions, which transform consistently under these transformations\. AfterLLlayers,๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limitsoutputs๐ฌi,Lp\\mathbf\{s\}^\{p\}\_\{i,L\}andโ„‹i,Lp\\mathcal\{H\}^\{p\}\_\{i,L\}as the final pocket atom embeddings, that is,

๐ฌip=๐ฌi,Lp,โ„‹ip=โ„‹i,Lp\.\\mathbf\{s\}^\{p\}\_\{i\}=\\mathbf\{s\}^\{p\}\_\{i,L\},\\mathcal\{H\}^\{p\}\_\{i\}=\\mathcal\{H\}^\{p\}\_\{i,L\}\.\(2\)
Specifically, at thell\-th layer,๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limitsupdates the embeddings๐ฌi,lp\\mathbf\{s\}^\{p\}\_\{i,l\}andโ„‹i,lp\\mathcal\{H\}^\{p\}\_\{i,l\}for atom๐–บip\\mathsf\{a\}^\{p\}\_\{i\}by aggregating information from its neighboring atoms as follows:

\(๐ฌi,lp,โ„‹i,lp\)\\displaystyle\(\\mathbf\{s\}^\{p\}\_\{i,l\},\\mathcal\{H\}^\{p\}\_\{i,l\}\)=GVPโ€‹\(๐กi,lp,๐’ดi,lp\),where\\displaystyle=\\text\{GVP\}\(\\mathbf\{h\}^\{p\}\_\{i,l\},\\mathcal\{Y\}^\{p\}\_\{i,l\}\),\\text\{ where \}\(3\)๐กi,lp\\displaystyle\\mathbf\{h\}^\{p\}\_\{i,l\}=\[๐ฏip,๐ฌi,lโˆ’1p,โˆ‘๐–บjpโˆˆ๐\(๐–บip\|๐’ซ\)ejโ€‹i,l๐ฆjโ€‹i,lp\],\\displaystyle=\\biggr\[\\mathbf\{v\}^\{p\}\_\{i\},\\mathbf\{s\}^\{p\}\_\{i,l\-1\},\\sum\_\{\\mathsf\{a\}^\{p\}\_\{j\}\\in\{\\mbox\{$\\mathop\{\\mathbf\{N\}\}\\limits$\}\(\\mathsf\{a\}^\{p\}\_\{i\}\|\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\)\}\}e\_\{ji,l\}\\mathbf\{m\}^\{p\}\_\{ji,l\}\\biggr\],๐’ดi,lp\\displaystyle\\mathcal\{Y\}^\{p\}\_\{i,l\}=\[๐ฑip,โ„‹i,lโˆ’1p,โˆ‘๐–บjpโˆˆ๐\(๐–บip\|๐’ซ\)ejโ€‹i,lโ€‹Mjโ€‹i,lp\],\\displaystyle=\\biggl\[\\mathbf\{x\}^\{p\}\_\{i\},\\mathcal\{H\}^\{p\}\_\{i,l\-1\},\\sum\_\{\\mathsf\{a\}^\{p\}\_\{j\}\\in\{\\mbox\{$\\mathop\{\\mathbf\{N\}\}\\limits$\}\(\\mathsf\{a\}^\{p\}\_\{i\}\|\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\)\}\}e\_\{ji,l\}M^\{p\}\_\{ji,l\}\\biggr\],whereGVPโ€‹\(โ‹…\)\\text\{GVP\}\(\\cdot\)models interactions between๐กi,lp\\mathbf\{h\}^\{p\}\_\{i,l\}and๐’ดi,lp\\mathcal\{Y\}^\{p\}\_\{i,l\}\. Intuitively,๐กi,lp\\mathbf\{h\}^\{p\}\_\{i,l\}aggregates the chemical characteristics of๐–บip\\mathsf\{a\}^\{p\}\_\{i\}, and๐’ดi,lp\\mathcal\{Y\}^\{p\}\_\{i,l\}aggregates its geometric features\.๐\(๐–บip\|๐’ซ\)\\mbox\{$\\mathop\{\\mathbf\{N\}\}\\limits$\}\(\\mathsf\{a\}^\{p\}\_\{i\}\|\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\)denotes the set of thekk\-nearest pocket atoms to๐–บip\\mathsf\{a\}^\{p\}\_\{i\}, selected based on proximity among atoms in๐’ซ\\mathop\{\\mathcal\{P\}\}\\limits;๐ฆjโ€‹i,lp\\mathbf\{m\}^\{p\}\_\{ji,l\}andMjโ€‹i,lpM^\{p\}\_\{ji,l\}correspond to the scalar and vector message embeddings, respectively, which propagate information from atom๐–บjpโˆˆ๐\(๐–บip\|๐’ซ\)\\mathsf\{a\}^\{p\}\_\{j\}\\in\\mbox\{$\\mathop\{\\mathbf\{N\}\}\\limits$\}\(\\mathsf\{a\}^\{p\}\_\{i\}\|\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\)to๐–บip\\mathsf\{a\}^\{p\}\_\{i\}; andejโ€‹i,le\_\{ji,l\}denotes the attention weight between neighbor atoms๐–บjp\\mathsf\{a\}^\{p\}\_\{j\}and๐–บip\\mathsf\{a\}^\{p\}\_\{i\}\. The message embeddings๐ฆjโ€‹i,lp\\mathbf\{m\}^\{p\}\_\{ji,l\}andMjโ€‹i,lpM^\{p\}\_\{ji,l\}are calculated as follows:

\(๐ฆjโ€‹i,lp,Mjโ€‹i,lp\)\\displaystyle\(\\mathbf\{m\}^\{p\}\_\{ji,l\},M^\{p\}\_\{ji,l\}\)=GVPโ€‹\(๐ฆ^jโ€‹i,lp,M^jโ€‹i,lp\),where\\displaystyle=\\text\{GVP\}\(\\hat\{\\mathbf\{m\}\}^\{p\}\_\{ji,l\},\\hat\{M\}^\{p\}\_\{ji,l\}\),\\text\{ where \}\(4\)๐ฆ^jโ€‹i,lp\\displaystyle\\hat\{\\mathbf\{m\}\}^\{p\}\_\{ji,l\}=\[๐ฌj,lโˆ’1p,dโ€‹\(๐–บjp,๐–บip\)\],\\displaystyle=\[\\mathbf\{s\}^\{p\}\_\{j,l\-1\},\{d\(\\mathsf\{a\}^\{p\}\_\{j\},\\mathsf\{a\}^\{p\}\_\{i\}\)\}\],M^jโ€‹i,lp\\displaystyle\\hat\{M\}^\{p\}\_\{ji,l\}=\[โ„‹j,lโˆ’1p,๐ฑjpโˆ’๐ฑip\],\\displaystyle=\[\\mathcal\{H\}^\{p\}\_\{j,l\-1\},\{\\mathbf\{x\}^\{p\}\_\{j\}\-\\mathbf\{x\}^\{p\}\_\{i\}\]\},wheredโ€‹\(๐–บjp,๐–บip\)d\(\\mathsf\{a\}^\{p\}\_\{j\},\\mathsf\{a\}^\{p\}\_\{i\}\)represents the Euclidean distance between the positions of pocket atoms๐–บjp\\mathsf\{a\}^\{p\}\_\{j\}and๐–บip\\mathsf\{a\}^\{p\}\_\{i\}\. The attention weightsejโ€‹i,le\_\{ji,l\}are calculated as follows:

ejโ€‹i,l\\displaystyle e\_\{ji,l\}=expโก\(Qi,lโ€‹Kjโ€‹i,l\)โˆ‘๐–บkpโˆˆ๐\(๐–บip\|๐’ซ\)expโก\(Qi,lโ€‹Kkโ€‹i,l\),where\\displaystyle=\\frac\{\\exp\(Q\_\{i,l\}K\_\{ji,l\}\)\}\{\\sum\_\{\\mathsf\{a\}^\{p\}\_\{k\}\\in\{\\mbox\{$\\mathop\{\\mathbf\{N\}\}\\limits$\}\(\\mathsf\{a\}^\{p\}\_\{i\}\|\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\)\}\}\\exp\(Q\_\{i,l\}K\_\{ki,l\}\)\},\\text\{ where \}\(5\)Qi,l\\displaystyle Q\_\{i,l\}=MLPโ€‹\(\[๐ฌi,lโˆ’1p,โ€–โ„‹i,lโˆ’1pโ€–2\]\),\\displaystyle=\\text\{MLP\}\(\{\[\\mathbf\{s\}^\{p\}\_\{i,l\-1\},\\\|\\mathcal\{H\}^\{p\}\_\{i,l\-1\}\\\|\_\{2\}\]\}\),Kjโ€‹i,l\\displaystyle K\_\{ji,l\}=MLPโ€‹\(\[๐ฆjโ€‹i,lp,โ€–Mjโ€‹i,lpโ€–2\]\),\\displaystyle=\\text\{MLP\}\(\[\\mathbf\{m\}^\{p\}\_\{ji,l\},\\\|M^\{p\}\_\{ji,l\}\\\|\_\{2\}\]\),whereโˆฅโ‹…โˆฅ2\\\|\\cdot\\\|\_\{2\}denotes the Euclidean \(โ„“2\\ell\_\{2\}\) norm of a 2D vector\. Intuitively, the attention weightejโ€‹i,le\_\{ji,l\}measures how much neighbor atom๐–บjp\\mathsf\{a\}^\{p\}\_\{j\}should influence the update of๐–บip\\mathsf\{a\}^\{p\}\_\{i\}based on both chemical and spatial information encoded inQi,jQ\_\{i,j\}andKjโ€‹i,lK\_\{ji,l\}, allowing the model to pay more attention to more informative neighbors\. Here,Qi,lQ\_\{i,l\}represents the query of atom๐–บip\\mathsf\{a\}^\{p\}\_\{i\}, summarizing the features of atom๐–บip\\mathsf\{a\}^\{p\}\_\{i\}and geometry surrounding it, whileKjโ€‹i,lK\_\{ji,l\}represents the key of neighbor๐–บjp\\mathsf\{a\}^\{p\}\_\{j\}, capturing its message\-level features that combine chemical and spatial information\.

##### Pocket Residue Embeddings

๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limitsalso learns residue embeddings, denoted as๐ฌir\\mathbf\{s\}^\{r\}\_\{i\}andโ„‹ir\\mathcal\{H\}^\{r\}\_\{i\}, respectively, using a graph neural network \(GNN\) similar to the one described above \(Equations[3](https://arxiv.org/html/2607.12349#Sx6.E3)to[5](https://arxiv.org/html/2607.12349#Sx6.E5)\) over residues\. Incorporating this residue information allows๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limitsto capture important physicochemical properties at the residue level, such as polarity, charge, and hydrophobicity, that are critical for determining protein\-ligand interactions\[gilson2007calculation\]\.

##### Pocket Embeddings

๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limitsintegrates pocket atom embeddings and pocket residue embeddings into a single, enriched pocket embedding, comprehensively capturing the multi\-granular spatial and physicochemical features of the protein pockets\. Particularly,๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limitscombines each pocket atom embedding and the embedding of its belonging residue through a gating mechanism, which balances the two embeddings by adaptively weighting their contributions\. Thus, the atom scalar embedding for๐–บip\\mathsf\{a\}^\{p\}\_\{i\}is updated as follows:

๐ฌip\\displaystyle\\mathbf\{s\}^\{p\}\_\{i\}=gisโ€‹๐ฌip\+\(1โˆ’gis\)โ€‹๐ฌjr,๐—‹jp=๐—‹pโ€‹\(๐–บip\),where\\displaystyle=g^\{s\}\_\{i\}\\mathbf\{s\}^\{p\}\_\{i\}\+\(1\-g^\{s\}\_\{i\}\)\\mathbf\{s\}^\{r\}\_\{j\},~~\\mathsf\{r\}^\{p\}\_\{j\}=\\mathsf\{r\}^\{p\}\(\\mathsf\{a\}^\{p\}\_\{i\}\),\\text\{ where \}\(6\)gis\\displaystyle g^\{s\}\_\{i\}=ฯƒ\(MLP\(\[๐ฌip,๐ฌjr\]\),\\displaystyle=\\sigma\(\\text\{MLP\}\(\[\\mathbf\{s\}^\{p\}\_\{i\},\\mathbf\{s\}^\{r\}\_\{j\}\]\),wheregisg^\{s\}\_\{i\}denotes the gating weight, learned from the scalar embeddings of the pocket atom๐–บig\\mathsf\{a\}^\{g\}\_\{i\}and its corresponding residue๐—‹pโ€‹\(๐–บip\)\\mathsf\{r\}^\{p\}\(\\mathsf\{a\}^\{p\}\_\{i\}\)\. Similarly, the atom vector embedding for๐–บip\\mathsf\{a\}^\{p\}\_\{i\}is updated as follows:

โ„‹ip\\displaystyle\\mathcal\{H\}^\{p\}\_\{i\}=gihโ€‹โ„‹ip\+\(1โˆ’gih\)โ€‹โ„‹jr,๐—‹jp=๐—‹pโ€‹\(๐–บip\),where\\displaystyle=g^\{h\}\_\{i\}\\mathcal\{H\}^\{p\}\_\{i\}\+\(1\-g^\{h\}\_\{i\}\)\\mathcal\{H\}^\{r\}\_\{j\},~~\\mathsf\{r\}^\{p\}\_\{j\}=\\mathsf\{r\}^\{p\}\(\\mathsf\{a\}^\{p\}\_\{i\}\),\\text\{ where \}\(7\)gih\\displaystyle g^\{h\}\_\{i\}=ฯƒโ€‹\(MLPโ€‹\(\[โ€–โ„‹ipโ€–2,โ€–โ„‹jrโ€–2\]\)\)\.\\displaystyle=\{\\sigma\(\\text\{MLP\}\(\[\\\|\\mathcal\{H\}^\{p\}\_\{i\}\\\|\_\{2\},\\\|\\mathcal\{H\}^\{r\}\_\{j\}\\\|\_\{2\}\]\)\)\}\.This gating mechanism enables the model to dynamically integrate atom\-level and residue\-level information, leading to expressive pocket representations over pocket atoms\.

#### Pocket Decoder \(๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–ฝ๐–พ๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{dec\}\}\\limits\)

To train๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits, we introduce a pocket decoder, denoted as๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–ฝ๐–พ๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{dec\}\}\\limits, to facilitate self\-supervised training on a contrived noise prediction task, inspired byUni\-Mol\[zhou2023uni\]\. To create supervision signals, we corrupt each pocket by applying two types of perturbations:\(1\)atom position perturbation โ€“ uniform noises are added to the positions of a randomly selected subset of atoms, and\(2\)atom type masking โ€“ the same subset of pocket atoms as in the position perturbing step has their atom types masked off\. The corrupted pocket is encoded to embeddings by๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits, as described in section Pocket Encoder \(๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits\)\.

๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–ฝ๐–พ๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{dec\}\}\\limitslearns to recover the original atom positions and types from the embeddings of noisy pockets provided by๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits\. Specifically,๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–ฝ๐–พ๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{dec\}\}\\limitspredicts the normal noise๐ง^ip\\hat\{\\mathbf\{n\}\}^\{p\}\_\{i\}added to the position of a pocket atom๐–บip\\mathsf\{a\}^\{p\}\_\{i\}using the following formulation:

๐ง^ip\\displaystyle\\hat\{\\mathbf\{n\}\}^\{p\}\_\{i\}=โˆ‘๐–บjpโˆˆ๐\(๐–บip\|๐’ซ\)cjโ€‹iโ€‹\(๐ฑjpโˆ’๐ฑip\)Ni,where\\displaystyle=\\sum\_\{\\mathsf\{a\}^\{p\}\_\{j\}\\in\{\\mbox\{$\\mathop\{\\mathbf\{N\}\}\\limits$\}\(\\mathsf\{a\}^\{p\}\_\{i\}\|\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\)\}\}\\frac\{c\_\{ji\}\(\{\\mathbf\{x\}^\{p\}\_\{j\}\-\\mathbf\{x\}^\{p\}\_\{i\}\}\)\}\{N\_\{i\}\},\\text\{ where \}\(8\)cjโ€‹i\\displaystyle\\quad c\_\{ji\}=MLPโ€‹\(\[๐ฌip,๐ฌjp,dโ€‹\(๐–บjp,๐–บip\),โ€–โ„‹jpโˆ’โ„‹ipโ€–2\]\),\\displaystyle=\\text\{MLP\}\(\[\\mathbf\{s\}^\{p\}\_\{i\},\\mathbf\{s\}^\{p\}\_\{j\},\{d\(\\mathsf\{a\}^\{p\}\_\{j\},\\mathsf\{a\}^\{p\}\_\{i\}\)\},\\\|\\mathcal\{H\}^\{p\}\_\{j\}\-\\mathcal\{H\}^\{p\}\_\{i\}\\\|\_\{2\}\]\),whereNi:=\|๐\(๐–บip\|๐’ซ\)\|N\_\{i\}:=\|\\mbox\{$\\mathop\{\\mathbf\{N\}\}\\limits$\}\(\\mathsf\{a\}^\{p\}\_\{i\}\|\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\)\|, which denotes the number of neighboring pocket atoms of๐–บip\\mathsf\{a\}^\{p\}\_\{i\}\. Thus,๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–ฝ๐–พ๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{dec\}\}\\limitsestimates the positional noise of each atom๐–บip\\mathsf\{a\}^\{p\}\_\{i\}by aggregating the relative position vectors from its neighboring atoms\. The weighting coefficientcjโ€‹ic\_\{ji\}from each neighboring atom๐–บjp\\mathsf\{a\}^\{p\}\_\{j\}is computed from the embeddings๐ฌp\\mathbf\{s\}^\{p\}andโ„‹p\\mathcal\{H\}^\{p\}produced by๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits\. This formulation exploits the local geometric structure of the pocket, forcing๐ฌp\\mathbf\{s\}^\{p\}andโ„‹p\\mathcal\{H\}^\{p\}to encode spatial relationships among neighboring atoms and reflect local geometric information\. We predict the added positional noise rather than the clean positions here since the noise follows a simpler and more tractable distribution, which stabilizes training and is easier to learn\.

For masked atom types,๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–ฝ๐–พ๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{dec\}\}\\limitspredicts a probability vector over all possible atom types for each masked atom๐–บip\\mathsf\{a\}^\{p\}\_\{i\}, denoted as๐ฏ^ip\\hat\{\\mathbf\{v\}\}^\{p\}\_\{i\}, by a multi\-layer perceptron \(MLP\):

๐ฏ^ip=softmaxโ€‹\(MLPโ€‹\(๐ฌip\)\)\.\\hat\{\\mathbf\{v\}\}^\{p\}\_\{i\}=\\text\{softmax\}\(\\text\{MLP\}\(\\mathbf\{s\}^\{p\}\_\{i\}\)\)\.\(9\)This formulation enables๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–ฝ๐–พ๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{dec\}\}\\limitsto infer atom types based on the learned pocket embedding๐ฌp\\mathbf\{s\}^\{p\}\. By designing๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–ฝ๐–พ๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{dec\}\}\\limitsto reconstruct atom types and positions, we guide๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limitsto learn pocket embeddings that capture both atomic characteristics and local geometry for accurate recovery and refinement\.

#### ๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\\mathop\{\\mathsf\{msPRL\}\}\\limitsPretraining

To pretrain๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\\mathop\{\\mathsf\{msPRL\}\}\\limits, followingUni\-Molโ€™s denoising pretraining\[zhou2023uni\], we randomly corrupt 20% of the atoms in pocket๐’ซ\\mathop\{\\mathcal\{P\}\}\\limits\. The original positions and atom types of the noisy pocket๐’ซ\\mathop\{\\mathcal\{P\}\}\\limitsare reconstructed using๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–ฝ๐–พ๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{dec\}\}\\limitsby minimizing the loss function:

โ„’p:=1\|โ„ณ\|โ€‹โˆ‘๐–บipโˆˆ๐’ซ๐•€โ€‹\(๐–บipโˆˆโ„ณ\)โ€‹\(โ€–๐ง^ipโˆ’๐งipโ€–2\+๐–งโ€‹\(๐ฏ^ip,๐ฏip\)\),\\mathcal\{L\}\_\{p\}:=\\frac\{1\}\{\|\\mathcal\{M\}\|\}\\sum\_\{\\mathsf\{a\}^\{p\}\_\{i\}\\in\{\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\}\}\\mathbb\{I\}\(\\mathsf\{a\}^\{p\}\_\{i\}\\in\\mathcal\{M\}\)\\,\\Big\(\\,\\\|\\hat\{\\mathbf\{n\}\}^\{p\}\_\{i\}\-\\mathbf\{n\}^\{p\}\_\{i\}\\\|^\{2\}\\;\+\\;\\mathsf\{H\}\(\\hat\{\\mathbf\{v\}\}^\{p\}\_\{i\},\\mathbf\{v\}^\{p\}\_\{i\}\)\\Big\),\(10\)whereโ„ณ\\mathcal\{M\}is the set of perturbed atoms in๐’ซ\\mathop\{\\mathcal\{P\}\}\\limits,๐•€โ€‹\(โ‹…\)\\mathbb\{I\}\(\\cdot\)denotes the indicator function,๐–งโ€‹\(โ‹…,โ‹…\)\\mathsf\{H\}\(\\cdot,\\cdot\)represents the cross\-entropy loss, and๐ง^ip\\hat\{\\mathbf\{n\}\}^\{p\}\_\{i\}and๐งip\\mathbf\{n\}^\{p\}\_\{i\}denote the predicted noise from the decoder and the ground\-truth noise from the corruption step, respectively\. The overall pretraining process is summarized in Algorithm[1](https://arxiv.org/html/2607.12349#alg1)\.

Algorithm 1๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\\mathop\{\\mathsf\{msPRL\}\}\\limitsRequired Input: all๐’ซ\\mathop\{\\mathcal\{P\}\}\\limitsin Training data of๐–ข๐–ฃ\\mathop\{\\mathsf\{CD\}\}\\limits

1:whilenot convergeddo

2:Sample a batch

โ„ฌ\\mathcal\{B\}of๐’ซ\\mathop\{\\mathcal\{P\}\}\\limits

3:

Lossโ†0\\text\{Loss\}\\leftarrow 0
4:for

๐’ซโˆˆโ„ฌ\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\\in\\mathcal\{B\}do

5:

๐ฑp,๐ฏp,๐ฑr,๐ฏr\\mathbf\{x\}^\{p\},\\mathbf\{v\}^\{p\},\\mathbf\{x\}^\{r\},\\mathbf\{v\}^\{r\}=

addNoiseโ€‹\(๐’ซ\)\\text\{addNoise\}\(\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\)โŠณ\\trianglerightperturb pocket structure

6:

๐ฌp,โ„‹p=๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\(๐ฑp,๐ฏp,๐ฑr,๐ฏr\)\\mathbf\{s\}^\{p\},\\mathcal\{H\}^\{p\}=\\mbox\{$\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits$\}\(\\mathbf\{x\}^\{p\},\\mathbf\{v\}^\{p\},\\mathbf\{x\}^\{r\},\\mathbf\{v\}^\{r\}\)โŠณ\\trianglerightencode the positions and features into latent embeddings

7:

๐’ซ^=๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–ฝ๐–พ๐–ผ\(๐ฌp,โ„‹p,๐ฑp\)\\hat\{\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\}=\\mbox\{$\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{dec\}\}\\limits$\}\(\\mathbf\{s\}^\{p\},\\mathcal\{H\}^\{p\},\\mathbf\{x\}^\{p\}\)โŠณ\\trianglerightreconstruct pocket by decoding the embeddings

8:

Lossโ†Loss\+โ„’pโ€‹\(๐’ซ^,๐’ซ\)\\text\{Loss\}\\leftarrow\\text\{Loss\}\+\\mathcal\{L\}\_\{p\}\(\\hat\{\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\},\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\)โŠณ\\trianglerightaccumulate reconstruction loss \(Equation[10](https://arxiv.org/html/2607.12349#Sx6.E10)\)

9:endfor

10:

Lossโ†1\|โ„ฌ\|โ€‹Loss\\text\{Loss\}\\leftarrow\\frac\{1\}\{\|\\mathcal\{B\}\|\}\\text\{Loss\}
11:

Loss\.backward\(๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ,๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–ฝ๐–พ๐–ผ\)\\text\{Loss\}\.\\text\{backward\(\\mbox\{$\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits$\},\\;\\mbox\{$\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{dec\}\}\\limits$\}\)\}โŠณ\\trianglerightEquation[10](https://arxiv.org/html/2607.12349#Sx6.E10)

12:Update๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits,๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–ฝ๐–พ๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{dec\}\}\\limits

13:endwhile

14:return๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits

### Pocket\-conditioned Ligand Generation via Conditional Diffusion \(๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limits\)

We present a new conditional diffusion model,๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limits, for the generation of three\-dimensional \(3D\) ligands that bind to specific, given protein pockets, referred to as condition pockets\. The proposed model generates realistic 3D ligands with high binding affinities by:\(1\)incorporating learned pocket structures and interactions within the pocket from๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits,\(2\)modeling atomic interactions within the ligand, and\(3\)learning pocket\-ligand interactions\. The subsequent sections provide a detailed description of the diffusion process \(Section Diffusion Process\), the training methodology of๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limits\(Section Model Training\), and the pocket\-conditioned ligand generator \(Section Pocket\-conditioned Ligand Generator \(๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limits\)\)\.

Following the framework of denoising diffusion probabilistic models\[ho2020ddpm\],๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitscomprises a forward diffusion process that progressively adds noise to the atomic positions\{๐ฑig\}\\\{\\mathbf\{x\}^\{g\}\_\{i\}\\\}and features\{๐ฏig\}\\\{\\mathbf\{v\}^\{g\}\_\{i\}\\\}of ligands, and a reverse generative process that is trained to denoise and generate new ligands during inference\. During training, the model learns to invert the forward corruption process by iteratively denoising noisy ligands\. During inference,๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsfirst samples noisy ligand atom positions and features at stepTTfrom predefined simple distributions and then reconstructs realistic 3D ligand structures by iteratively removing noise untilt=0t=0\. At each reverse steptt, given a noisy structure\{\(๐ฑi,tg,๐ฏi,tg\)\}\\\{\(\\mathbf\{x\}^\{g\}\_\{i,t\},\\mathbf\{v\}^\{g\}\_\{i,t\}\)\\\},๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsemploys a GNN\-based ligand generator that uses fixed pocket embeddings from๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limitstogether with pocket atom/residue positions and types to predict the noise\-free atom positions\{๐ฑi,0g\}\\\{\\mathbf\{x\}^\{g\}\_\{i,0\}\\\}and features\{๐ฏi,0g\}\\\{\\mathbf\{v\}^\{g\}\_\{i,0\}\\\}of the ligand atoms\. Unlike other diffusion\-based SBDD methods\[guan2023targetdiff,guan2023decompdiff,huang2024protein,gu2024aligning\],๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limits\(1\)leverages pocket positions and types exclusively to model pocket\-ligand interactions during denoising, while the pocket structure remains captured in fixed embeddings produced by๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits\. The fixed pocket representation prevents embedding drifts and re\-encoding artifacts across๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsโ€™s denoising process, so the ligand generator can learn pocket\-ligand interactions efficiently and with greater coherence across the generation process\. Like prior SBDD works\[guan2023targetdiff,guan2023decompdiff,huang2024protein,gu2024aligning\],๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limits\(2\)models the atomic interactions between the ligand and pocket\. In addition to atom\-level modeling,๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsfurther integrates residue\-level interactions\. This multi\-scale featured design provides the ligand with a hierarchical view of the pocket, enabling a comprehensive understanding of pocket\-ligand interactions\.

#### Forward Diffusion Process \(๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹โ€‹\-โ€‹๐–ฟ๐—ˆ๐—‹๐—๐–บ๐—‹๐–ฝ\\mathop\{\\mathsf\{conDitar\}\\text\{\-\}\\mathsf\{forward\}\}\\limits\)

Following the previous work\[guan2023targetdiff\], each ligand atom position is progressively noised according to a Gaussian transition in the forward process\[ho2020ddpm\]:

qโ€‹\(๐ฑi,tgโˆฃ๐ฑi,tโˆ’1g\)=๐’ฉโ€‹\(๐ฑi,tg;1โˆ’ฮฒt๐ฑโ€‹๐ฑi,tโˆ’1g,ฮฒt๐ฑโ€‹๐ˆ\),q\(\\mathbf\{x\}^\{g\}\_\{i,t\}\\mid\\mathbf\{x\}^\{g\}\_\{i,t\-1\}\)=\\mathcal\{N\}\\\!\\left\(\\mathbf\{x\}^\{g\}\_\{i,t\};\\sqrt\{1\-\\beta^\{\\mathbf\{x\}\}\_\{t\}\}\\,\\mathbf\{x\}^\{g\}\_\{i,t\-1\},\\beta^\{\\mathbf\{x\}\}\_\{t\}\\mathbf\{I\}\\right\),\(11\)whereฮฒt๐ฑ\\beta^\{\\mathbf\{x\}\}\_\{t\}is a predefined variance schedule that controls the noise magnitude added at steptt;๐ˆ\\mathbf\{I\}denotes the identity matrix\. OverTTsteps, this produces a Markov chain that gradually transforms the clean coordinates into isotropic Gaussian noise\. Similarly, atom types are corrupted through a categorical diffusion process\. At each steptt, the current atom type is retained with probability1โˆ’ฮฒt๐ฏ1\-\\beta^\{\\mathbf\{v\}\}\_\{t\}, while the remaining probability mass is uniformly distributed among the other possible types:

qโ€‹\(๐ฏi,tgโˆฃ๐ฏi,tโˆ’1g\)=๐’žโ€‹\(๐ฏi,tg;\(1โˆ’ฮฒt๐ฏ\)โ€‹๐ฏi,tโˆ’1g\+ฮฒt๐ฏdaโ€‹๐Ÿ\),q\(\\mathbf\{v\}^\{g\}\_\{i,t\}\\mid\\mathbf\{v\}^\{g\}\_\{i,t\-1\}\)=\\mathcal\{C\}\\\!\\Big\(\\mathbf\{v\}^\{g\}\_\{i,t\};\\,\(1\-\\beta^\{\\mathbf\{v\}\}\_\{t\}\)\\,\{\\mathbf\{v\}^\{g\}\_\{i,t\-1\}\}\\;\+\\;\\frac\{\\beta^\{\\mathbf\{v\}\}\_\{t\}\}\{d\_\{a\}\}\\mathbf\{1\}\\Big\),\(12\)whereฮฒt๐ฏ\\beta^\{\\mathbf\{v\}\}\_\{t\}controls the level of corruption;๐Ÿ\\mathbf\{1\}denotes the all\-ones vector indicating uniform probability over alldad\_\{a\}atom types\.

#### Reverse Generative Process \(๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹โ€‹\-โ€‹๐—‹๐–พ๐—๐–พ๐—‹๐—Œ๐–พ\\mathop\{\\mathsf\{conDitar\}\\text\{\-\}\\mathsf\{reverse\}\}\\limits\)

In the reverse process,๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitslearns to reverse the forward process and generates ligands by denoising from\{\(๐ฑi,tg,๐ฏi,tg\)\}\\\{\(\\mathbf\{x\}^\{g\}\_\{i,t\},\\mathbf\{v\}^\{g\}\_\{i,t\}\)\\\}to\{\(๐ฑi,tโˆ’1g,๐ฏi,tโˆ’1g\)\}\\\{\(\\mathbf\{x\}^\{g\}\_\{i,t\-1\},\\mathbf\{v\}^\{g\}\_\{i,t\-1\}\)\\\}\. Following Ho*et al\.*\[ho2020ddpm\],๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsapproximates the probability of \(๐ฑi,tโˆ’1g\\mathbf\{x\}^\{g\}\_\{i,t\-1\},๐ฏi,tโˆ’1g\\mathbf\{v\}^\{g\}\_\{i,t\-1\}\) denoised from \(๐ฑi,tg\\mathbf\{x\}^\{g\}\_\{i,t\},๐ฏi,tg\\mathbf\{v\}^\{g\}\_\{i,t\}\) using the posteriorpโ€‹\(๐ฑi,tโˆ’1g\|๐ฑi,tg,๐ฑi,0g\)p\(\\mathbf\{x\}^\{g\}\_\{i,t\-1\}\|\\mathbf\{x\}^\{g\}\_\{i,t\},\\mathbf\{x\}^\{g\}\_\{i,0\}\)andpโ€‹\(๐ฏi,tโˆ’1g\|๐ฏi,tg,๐ฏi,0g\)p\(\\mathbf\{v\}^\{g\}\_\{i,t\-1\}\|\\mathbf\{v\}^\{g\}\_\{i,t\},\\mathbf\{v\}^\{g\}\_\{i,0\}\)\. Given that๐ฑi,0g\\mathbf\{x\}^\{g\}\_\{i,0\}and๐ฏi,0g\\mathbf\{v\}^\{g\}\_\{i,0\}are unknown in the backward process, at each steptt,๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsuses a new pocket\-conditioned ligand generator,๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limits\(Pocket\-conditioned Ligand Generator \(๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limits\) section\), to estimate them as follows:

\{\(๐ฑ~i,0,tg,๐ฏ~i,0,tg\)\}=๐—‰๐–ผ๐–ซ๐–ฆ\(\{๐ฑi,tg\},\{๐ฏi,tg\},\{๐ฌjp\},\{โ„‹jp\}\),\\\{\(\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{i,0,t\},\\tilde\{\\mathbf\{v\}\}^\{g\}\_\{i,0,t\}\)\\\}=\\mbox\{$\\mathop\{\\mathsf\{pcLG\}\}\\limits$\}\(\\\{\\mathbf\{x\}^\{g\}\_\{i,t\}\\\},\\\{\\mathbf\{v\}^\{g\}\_\{i,t\}\\\},\\\{\\mathbf\{s\}^\{p\}\_\{j\}\\\},\\\{\\mathcal\{H\}^\{p\}\_\{j\}\\\}\),\(13\)where๐ฑ~i,0,tg\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{i,0,t\}and๐ฏ~i,0,tg\\tilde\{\\mathbf\{v\}\}^\{g\}\_\{i,0,t\}are the estimates of๐ฑi,0g\\mathbf\{x\}^\{g\}\_\{i,0\}and๐ฏi,0g\\mathbf\{v\}^\{g\}\_\{i,0\}at timesteptt, respectively\.\{๐ฌjp\}\\\{\\mathbf\{s\}^\{p\}\_\{j\}\\\}and\{โ„‹jp\}\\\{\\mathcal\{H\}^\{p\}\_\{j\}\\\}represent the pretrained scalar and vector pocket embeddings\)\. Using the estimates,๐ฑi,tโˆ’1g\\mathbf\{x\}^\{g\}\_\{i,t\-1\}can be sampled as:

pโ€‹\(๐ฑi,tโˆ’1gโˆฃ๐ฑi,tg\)\\displaystyle p\(\\mathbf\{x\}^\{g\}\_\{i,t\-1\}\\mid\\mathbf\{x\}^\{g\}\_\{i,t\}\)โ‰ˆqโ€‹\(๐ฑi,tโˆ’1gโˆฃ๐ฑi,tg,๐ฑ~i,0,tg\)=๐’ฉโ€‹\(๐ฑi,tโˆ’1gโˆฃฮผโ€‹\(๐ฑi,tg,๐ฑ~i,0,tg\),ฮฒ~t๐ฑโ€‹๐ˆ\),\\displaystyle\\approx q\(\\mathbf\{x\}^\{g\}\_\{i,t\-1\}\\mid\\mathbf\{x\}^\{g\}\_\{i,t\},\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{i,0,t\}\)=\\mathcal\{N\}\(\\mathbf\{x\}^\{g\}\_\{i,t\-1\}\\mid\\mu\(\\mathbf\{x\}^\{g\}\_\{i,t\},\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{i,0,t\}\),\\tilde\{\\beta\}\_\{t\}^\{\\mathbf\{x\}\}\\mathbf\{I\}\),\(14\)whereฮผโ€‹\(๐ฑi,tg,๐ฑ~i,0,tg\)\\mu\(\\mathbf\{x\}^\{g\}\_\{i,t\},\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{i,0,t\}\)is the estimated mean of Gaussian transition distribution \(detailed in Appendix[C](https://arxiv.org/html/2607.12349#A3)\)\. Similarly,๐ฏi,tโˆ’1g\\mathbf\{v\}^\{g\}\_\{i,t\-1\}can be sampled as

pโ€‹\(๐ฏi,tโˆ’1gโˆฃ๐ฏi,tg\)โ‰ˆqโ€‹\(๐ฏi,tโˆ’1gโˆฃ๐ฏi,tg,๐ฏ~i,0,tg\)=๐’žโ€‹\(๐ฏi,tโˆ’1gโˆฃ๐œโ€‹\(๐ฏi,tg,๐ฏ~i,0,tg\)\),p\(\\mathbf\{v\}^\{g\}\_\{i,t\-1\}\\mid\\mathbf\{v\}^\{g\}\_\{i,t\}\)\\approx q\(\\mathbf\{v\}^\{g\}\_\{i,t\-1\}\\mid\\mathbf\{v\}^\{g\}\_\{i,t\},\\tilde\{\\mathbf\{v\}\}^\{g\}\_\{i,0,t\}\)=\\mathcal\{C\}\(\\mathbf\{v\}^\{g\}\_\{i,t\-1\}\\mid\\mathbf\{c\}\(\\mathbf\{v\}^\{g\}\_\{i,t\},\\tilde\{\\mathbf\{v\}\}^\{g\}\_\{i,0,t\}\)\),\(15\)where๐œโ€‹\(๐ฏi,tg,๐ฏ~i,0,tg\)\\mathbf\{c\}\\big\(\\mathbf\{v\}^\{g\}\_\{i,t\},\\tilde\{\\mathbf\{v\}\}^\{g\}\_\{i,0,t\}\\big\)denotes the estimated probability vector that parameterizes the categorical transition distribution \(detailed in Appendix[C](https://arxiv.org/html/2607.12349#A3)\)\. In practice, the denoising update for๐ฑi,tโˆ’1g\\mathbf\{x\}^\{g\}\_\{i,t\-1\}can be calculated as

๐ฑi,tโˆ’1g=ฮผโ€‹\(๐ฑi,tg,๐ฑ~i,0,tg\)\+ฮฒ~t๐ฑโ€‹๐ณi,tโˆ’1๐ฑ,\\mathbf\{x\}^\{g\}\_\{i,t\-1\}=\\mu\(\\mathbf\{x\}^\{g\}\_\{i,t\},\\tilde\{\\mathbf\{x\}\}\_\{i,0,t\}^\{g\}\)\+\\sqrt\{\\tilde\{\\beta\}^\{\\mathbf\{x\}\}\_\{t\}\}\\mathbf\{z\}^\{\\mathbf\{x\}\}\_\{i,t\-1\},\(16\)where๐ณi,tโˆ’1๐ฑโˆผ๐’ฉโ€‹\(๐ŸŽ,๐ˆ\)\\mathbf\{z\}^\{\\mathbf\{x\}\}\_\{i,t\-1\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)is a standard Gaussian noise sampled at timett, and the denoising update for๐ฏi,tโˆ’1g\\mathbf\{v\}^\{g\}\_\{i,t\-1\}can be calculated as

๐ฏi,tโˆ’1g=onehotโก\(argโกmaxโ„“โก\[logโก๐œโ€‹\(๐ฏi,tg,๐ฏ~i,0,tg\)โ„“\+๐ณi,tโˆ’1,โ„“๐šŸ\]\),\\mathbf\{v\}^\{g\}\_\{i,t\-1\}=\\operatorname\{onehot\}\\\!\\left\(\\arg\\max\_\{\\ell\}\\big\[\\,\\log\\mathbf\{c\}\(\\mathbf\{v\}^\{g\}\_\{i,t\},\\tilde\{\\mathbf\{v\}\}^\{g\}\_\{i,0,t\}\)\_\{\\ell\}\+\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{i,t\-1,\\ell\}\\,\\big\]\\right\),\(17\)where๐ณi,tโˆ’1,โ„“๐ฏโˆผGumbelโ€‹\(0,1\)\\mathbf\{z\}\_\{i,t\-1,\\ell\}^\{\\mathbf\{v\}\}\\sim\\text\{Gumbel\}\(0,1\)forโ„“=1,โ€ฆ,da\\ell=1,\\dots,d\_\{a\}\(dad\_\{a\}is the total number of atom types\) are independent Gumbel noises\. Here, we use the Gumbel\-max trick to sample from the discrete categorical distribution\. More details of the backward process are available in Appendix[C](https://arxiv.org/html/2607.12349#A3)\.

##### Determining Ligand Sizes

During the generative process, the ligand size \(i\.e\., the number of heavy atoms\) is unknown and must be determined before atom positions and types can be generated\. The ligand size is sampled based on the volume of the condition pocket, referred to as the pocket size\. The pocket size is determined by the region in proximity to the reference ligand\. Pocket atoms that are among the three nearest neighbors of the ligand atoms are selected, and the maximum pairwise distance among them defines the pocket size\. Pocket sizes are grouped into ordered intervals, each associated with an empirical distribution of ligand atom counts from the training data\. The ligand size is then sampled from the empirical distribution corresponding to the pocketโ€™s interval\.

#### ๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsModel Training

๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsis trained to recover original ligand atom positions๐ฑi,0g\\mathbf\{x\}^\{g\}\_\{i,0\}and features๐ฏi,0g\\mathbf\{v\}^\{g\}\_\{i,0\}, conditioned on the protein pocket\. Specifically,๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsuses the combination of the following three losses as its loss function\.

##### Atom Position Loss

This loss measures the Euclidean errors between the predicted positions๐ฑ~i,0,tg\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{i,0,t\}at stepttand the ground\-truth positions๐ฑi,0g\\mathbf\{x\}^\{g\}\_\{i,0\}as follows:

โ„’t๐ฑโ€‹\(๐’Ÿ\)\\displaystyle\\mathcal\{L\}\_\{t\}^\{\\mathbf\{x\}\}\(\)=wt๐ฑโˆ‘โˆ€๐–บigโˆˆ๐’ŸKL\(q\(๐ฑi,tโˆ’1g\|๐ฑi,tg,๐ฑi,0g\)\|\|q\(๐ฑi,tโˆ’1g\|๐ฑi,tg,๐ฑ~i,0,tg\)\)\\displaystyle=w\_\{t\}^\{\\mathbf\{x\}\}\\sum\_\{\\forall\\mathsf\{a\}^\{g\}\_\{i\}\\in\{\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\}\}\\text\{KL\}\(q\(\\mathbf\{x\}^\{g\}\_\{i,t\-1\}\|\\mathbf\{x\}^\{g\}\_\{i,t\},\\mathbf\{x\}^\{g\}\_\{i,0\}\)\|\|q\(\\mathbf\{x\}^\{g\}\_\{i,t\-1\}\|\\mathbf\{x\}^\{g\}\_\{i,t\},\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{i,0,t\}\)\)\(18\)=wt๐ฑโ€‹โˆ‘โˆ€๐–บigโˆˆ๐’Ÿโ€–๐ฑ~i,0,tgโˆ’๐ฑi,0gโ€–,\\displaystyle=w\_\{t\}^\{\\mathbf\{x\}\}\\sum\_\{\\forall\\mathsf\{a\}^\{g\}\_\{i\}\\in\{\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\}\}\\\|\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{i,0,t\}\-\\mathbf\{x\}^\{g\}\_\{i,0\}\\\|,wherewt๐ฑw\_\{t\}^\{\\mathbf\{x\}\}is a time\-dependent weight at steptt, calculated by the signal\-to\-noise ratio\[kingma2021variational\]\(detailed in Appendix[B](https://arxiv.org/html/2607.12349#A2)\), and KL is the Kullback\-Leibler divergence\[kullback1951information\]\. Asttincreases,ฮฑยฏt๐ฑ\\bar\{\\alpha\}\_\{t\}^\{\\mathbf\{x\}\}decreases monotonically, resulting in an increase in noise level and a corresponding decrease inwt๐ฑw\_\{t\}^\{\\mathbf\{x\}\}until it reaches the thresholdฮด\\delta\. As a result, it encourages the model to emphasize more accurate reconstruction of molecular structures when the data have low noise levels and contain sufficient signals\.

##### Atom Type Loss

This loss measures the KL divergence\[kullback1951information\]between the ground\-truth posteriorqโ€‹\(๐ฏi,tโˆ’1g\|๐ฏi,tg,๐ฏi,0g\)q\(\\mathbf\{v\}^\{g\}\_\{i,t\-1\}\|\\mathbf\{v\}^\{g\}\_\{i,t\},\\mathbf\{v\}^\{g\}\_\{i,0\}\)and its estimateqโ€‹\(๐ฏi,tโˆ’1g\|๐ฏtg,๐ฏ~i,0,tg\)q\(\\mathbf\{v\}^\{g\}\_\{i,t\-1\}\|\\mathbf\{v\}^\{g\}\_\{t\},\\tilde\{\\mathbf\{v\}\}^\{g\}\_\{i,0,t\}\)as follows:

โ„’t๐ฏโ€‹\(๐’Ÿ\)\\displaystyle\\mathcal\{L\}\_\{t\}^\{\\mathbf\{v\}\}\(\)=โˆ‘โˆ€๐–บigโˆˆ๐’ŸKL\(q\(๐ฏi,tโˆ’1g\|๐ฏi,tg,๐ฏi,0g\)\|\|q\(๐ฏi,tโˆ’1g\|๐ฏi,tg,๐ฏ~i,0,tg\)\)\\displaystyle=\\sum\_\{\\forall\\mathsf\{a\}^\{g\}\_\{i\}\\in\{\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\}\}\\text\{KL\}\(q\(\\mathbf\{v\}^\{g\}\_\{i,t\-1\}\|\\mathbf\{v\}^\{g\}\_\{i,t\},\\mathbf\{v\}^\{g\}\_\{i,0\}\)\|\|q\(\\mathbf\{v\}^\{g\}\_\{i,t\-1\}\|\\mathbf\{v\}^\{g\}\_\{i,t\},\\tilde\{\\mathbf\{v\}\}^\{g\}\_\{i,0,t\}\)\)\(19\)=โˆ‘โˆ€๐–บigโˆˆ๐’ŸKL\(๐œ\(๐ฏi,tg,๐ฏi,0g\)\|\|๐œ\(๐ฏi,tg,๐ฏ~i,0,tg\)\),\\displaystyle=\\sum\_\{\\forall\\mathsf\{a\}^\{g\}\_\{i\}\\in\{\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\}\}\\text\{KL\}\(\\mathbf\{c\}\(\\mathbf\{v\}^\{g\}\_\{i,t\},\\mathbf\{v\}^\{g\}\_\{i,0\}\)\|\|\\mathbf\{c\}\(\\mathbf\{v\}^\{g\}\_\{i,t\},\\tilde\{\\mathbf\{v\}\}^\{g\}\_\{i,0,t\}\)\),where๐œโ€‹\(๐ฏi,tg,๐ฏi,0g\)\\mathbf\{c\}\(\\mathbf\{v\}^\{g\}\_\{i,t\},\\mathbf\{v\}^\{g\}\_\{i,0\}\)is the categorical distribution of๐ฏi,tg\\mathbf\{v\}^\{g\}\_\{i,t\}and๐œฮ˜โ€‹\(๐ฏi,tg,๐ฏ~i,0,tg\)\\mathbf\{c\}\_\{\\Theta\}\(\\mathbf\{v\}^\{g\}\_\{i,t\},\\tilde\{\\mathbf\{v\}\}^\{g\}\_\{i,0,t\}\)is an estimate of๐œโ€‹\(๐ฏi,tg,๐ฏi,0g\)\\mathbf\{c\}\(\\mathbf\{v\}^\{g\}\_\{i,t\},\\mathbf\{v\}^\{g\}\_\{i,0\}\)\.

##### Bond Type Loss

Following the literature\[chen2025generating\], this loss is to help๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsbetter understand the relations among atoms\. By explicitly learning bonds, the model better enforces chemical validity and yields more realistic molecular structures\. To achieve this, at each time steptt, for eachll\-th layer of ligand generator๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limits\(Pocket\-conditioned Ligand Generator \(๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limits\) section\),๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsis optimized to accurately predict the bond types among pairs of ligand atoms\. Specifically, the bond type loss is defined as follows:

โ„’t,l๐š‹โ€‹\(๐’Ÿ\)=โˆ‘โˆ€๐–บigโˆˆ๐’Ÿโˆ‘โˆ€๐–บjgโˆˆ๐\(๐–บig\|๐’Ÿ\)๐–งโ€‹\(๐žiโ€‹j,t,l,๐›iโ€‹j\),\\mathcal\{L\}^\{\\mathtt\{b\}\}\_\{t,l\}\(\{\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\}\)=\\sum\_\{\\forall\\mathsf\{a\}^\{g\}\_\{i\}\\in\{\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\}\}~\\sum\_\{\\forall\\mathsf\{a\}^\{g\}\_\{j\}\\in\{\\mbox\{$\\mathop\{\\mathbf\{N\}\}\\limits$\}\(\\mathsf\{a\}^\{g\}\_\{i\}\|\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\)\}\}\\mathsf\{H\}\(\\mbox\{$\\mathop\{\\mathbf\{e\}\}\\limits$\}\_\{ij,t,l\},\\mbox\{$\\mathop\{\\mathbf\{b\}\}\\limits$\}\_\{ij\}\),\(20\)where๐–งโ€‹\(โ‹…\)\\mathsf\{H\}\(\\cdot\)denotescross\-entropyloss,๐žiโ€‹j,t,l\\mbox\{$\\mathop\{\\mathbf\{e\}\}\\limits$\}\_\{ij,t,l\}represents the bond type predictions \(Equation[27](https://arxiv.org/html/2607.12349#Sx6.E27)\) between atoms๐–บig\\mathsf\{a\}^\{g\}\_\{i\}and๐–บjg\\mathsf\{a\}^\{g\}\_\{j\}at thell\-th layer at time steptt;๐\(๐–บig\|๐’Ÿ\)\\mbox\{$\\mathop\{\\mathbf\{N\}\}\\limits$\}\(\\mathsf\{a\}^\{g\}\_\{i\}\|\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\)denotes thekk\-nearest ligand atoms of atom๐–บig\\mathsf\{a\}^\{g\}\_\{i\}in position๐ฑi,tg\\mathbf\{x\}^\{g\}\_\{i,t\}; and๐›iโ€‹j\\mbox\{$\\mathop\{\\mathbf\{b\}\}\\limits$\}\_\{ij\}denotes one\-hot vector indicating the ground\-truth bond type between๐–บig\\mathsf\{a\}^\{g\}\_\{i\}and๐–บjg\\mathsf\{a\}^\{g\}\_\{j\}\. The total bond type prediction loss is computed by aggregating losses across different layers as follows:

โ„’t๐›โ€‹\(๐’Ÿ\)=wt๐ฑLโˆ’1โ€‹โˆ‘l=1Lโˆ’1โ„’t,l๐›โ€‹\(๐’Ÿ\)\+wt๐ฑโ€‹โ„’t,L๐›โ€‹\(๐’Ÿ\),\\mathcal\{L\}^\{\\mathbf\{b\}\}\_\{t\}\(\{\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\}\)=\\frac\{w\_\{t\}^\{\\mathbf\{x\}\}\}\{L\-1\}\\sum\_\{l=1\}^\{L\-1\}\\mathcal\{L\}^\{\\mathbf\{b\}\}\_\{t,l\}\(\{\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\}\)\+w\_\{t\}^\{\\mathbf\{x\}\}\\mathcal\{L\}^\{\\mathbf\{b\}\}\_\{t,L\}\(\{\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\}\),\(21\)wherewt๐ฑw\_\{t\}^\{\\mathbf\{x\}\}is the same timestep weight used in the atom position loss;LLrepresents the total number of layers in thepocket\-conditionedligand generator๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limits\. Similar toโ„’t๐ฑโ€‹\(๐’Ÿ\)\\mathcal\{L\}^\{\\mathbf\{x\}\}\_\{t\}\(\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\)\(Equation[18](https://arxiv.org/html/2607.12349#Sx6.E18)\), the weightwt๐ฑw\_\{t\}^\{\\mathbf\{x\}\}encourages the model to focus more on accurately predicting bond types when the data provides sufficient signals, rather than being confused by major noises in the data\. Note that, similar to Jumper*et al\.*\[Jumper2021\], in Equation[21](https://arxiv.org/html/2607.12349#Sx6.E21),๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsassigns different weights to the final layer \(i\.e\.,ll=LL\) compared to the other layers, as we empirically find this design benefits the generation performance\.

##### Overall๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsloss

The overall loss function for๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsis defined as follows:

โ„’diff=๐”ผ๐’Ÿโˆผ๐™ณโ€‹๐”ผtโˆผ๐’ฐโ€‹\(1,T\)โ€‹\(โ„’t๐ฑโ€‹\(๐’Ÿ\)\+ฮพโ€‹โ„’t๐ฏโ€‹\(๐’Ÿ\)\+ฮถโ€‹โ„’t๐›โ€‹\(๐’Ÿ\)\),\\mathcal\{L\}\_\{\\text\{diff\}\}=\\mathbb\{E\}\_\{\{\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\\sim\\text\{$\\mathtt\{D\}$\}\}\}\\mathbb\{E\}\_\{t\\sim\\mathcal\{U\}\(1,T\)\}\\big\(\\mathcal\{L\}^\{\\mathbf\{x\}\}\_\{t\}\(\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\)\+\\xi\\,\\mathcal\{L\}^\{\\mathbf\{v\}\}\_\{t\}\(\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\)\+\\zeta\\,\\mathcal\{L\}^\{\\mathbf\{b\}\}\_\{t\}\(\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\)\\big\),\(22\)where๐™ณ\\mathtt\{D\}is the training binding complexes;T\{T\}is the total number of diffusion timesteps;๐’ฐโ€‹\(1,T\)\\mathcal\{U\}\(1,T\)represents uniform distribution over these timesteps;ฮพ\>0\\xi\>0andฮถ\>0\\zeta\>0are twohyper\-parametersthat balanceโ„’t๐ฑ\\mathcal\{L\}^\{\\mathbf\{x\}\}\_\{t\}\(๐’Ÿ\\mathop\{\\mathcal\{D\}\}\\limits\),โ„’t๐ฏ\\mathcal\{L\}^\{\\mathbf\{v\}\}\_\{t\}\(๐’Ÿ\\mathop\{\\mathcal\{D\}\}\\limits\) andโ„’t๐›โ€‹\(๐’Ÿ\)\\mathcal\{L\}^\{\\mathbf\{b\}\}\_\{t\}\(\{\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\}\)\.

### Pocket\-conditioned Ligand Generator \(๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limits\)

๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsemploys a pocket\-conditioned ligand generator,๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limits, to denoise noisy ligand atom positions๐ฑ~i,tg\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{i,t\}and features๐ฏ~i,tg\\tilde\{\\mathbf\{v\}\}^\{g\}\_\{i,t\}at each diffusion step\. At a high level, denoising involves predicting the added noises and recovering the original ligand structures by leveraging chemical information and spatial information, including pocket geometry encoded by pretrained๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits, ligand geometry, and the ligandโ€™s relative spatial arrangement within the pocket\. To guide this process,๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitslearns expressive ligand atom representations by modeling two types of interactions:\(1\)interactions within the ligand, which model relationships among neighboring atoms to capture each atomโ€™s local environment, and\(2\)pocket\-ligand interactions, which capture pocket circumstances to provide pocket\-informed ligand representations\. Through modeling these interactions,๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitscaptures both the ligandโ€™s structure and its external interaction patterns with the pocket\. Furthermore, by pretraining pocket representation,๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitsavoids modeling pocket structure and instead focuses on learning intra\-ligand and pocket\-ligand interactions\. When no ambiguity arises, we will eliminate subscriptttin the notations and use\(๐ฑig,๐ฏig\)\(\\mathbf\{x\}^\{g\}\_\{i\},\\mathbf\{v\}^\{g\}\_\{i\}\)for brevity\.

#### Ligand Representation Learning

This section describes how๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitslearns ligand representations \(embeddings with superscriptg\) that capture both internal molecular structure and interactions with the pocket\.๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitsemploys a multi\-layer GNN to iteratively update ligand atom embeddings based on intra\-ligand and pocket\-ligand interactions\. Specifically, at thell\-th layer of the GNN, for each ligand atom๐–บig\\mathsf\{a\}^\{g\}\_\{i\},๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitslearns an invariant scalar embedding๐ฌi,lgโˆˆโ„dร—1\\mathbf\{s\}^\{g\}\_\{i,l\}\\in\\mathbb\{R\}^\{d\\times 1\}to capture atom features and an equivariant vector embeddingโ„‹i,lgโˆˆโ„dร—3\\mathcal\{H\}^\{g\}\_\{i,l\}\\in\\mathbb\{R\}^\{d\\times 3\}for 3D structural features\. For๐–บig\\mathsf\{a\}^\{g\}\_\{i\},๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitsupdates๐ฌi,lg\\mathbf\{s\}^\{g\}\_\{i,l\}andโ„‹i,lg\\mathcal\{H\}^\{g\}\_\{i,l\}by aggregating information from neighboring atoms of๐–บig\\mathsf\{a\}^\{g\}\_\{i\}in the current noisy ligand and incorporating interactions with neighboring pocket atoms\. The embeddings๐ฌi,lg\\mathbf\{s\}^\{g\}\_\{i,l\}andโ„‹i,lg\\mathcal\{H\}^\{g\}\_\{i,l\}are updated as follows:

\(๐ฌi,lg,โ„‹i,lg\)=GVPโ€‹\(๐กi,l,๐’ดi,l\),where\(\\mathbf\{s\}^\{g\}\_\{i,l\},\\mathcal\{H\}^\{g\}\_\{i,l\}\)=\\text\{GVP\}\(\\mathbf\{h\}\_\{i,l\},\\mathcal\{Y\}\_\{i,l\}\),\\text\{ where \}\(23\)๐กi,l\\displaystyle\\mathbf\{h\}\_\{i,l\}=\[๐ฏig,๐ฌi,lโˆ’1gโŸfrom the ligand,๐ฌi,lโˆ’1c,๐ฌi,lโˆ’1a,โŸfrom the pocketโˆ‘๐–บjgโˆˆ๐\(๐–บig\|๐’Ÿ\)zjโ€‹i,l๐ฆjโ€‹i,lg\],\\displaystyle=\\biggr\[\\underbrace\{\\mathbf\{v\}^\{g\}\_\{i\},\\mathbf\{s\}^\{g\}\_\{i,l\-1\}\}\_\{\\mathclap\{\\text\{from the ligand\}\}\},~~~\\underbrace\{\\mathbf\{s\}^\{c\}\_\{i,l\-1\},\\mathbf\{s\}^\{a\}\_\{i,l\-1\},\}\_\{\\mathclap\{\\text\{from the pocket\}\}\}~\\sum\_\{\\mathsf\{a\}^\{g\}\_\{j\}\\in\{\\mbox\{$\\mathop\{\\mathbf\{N\}\}\\limits$\}\(\\mathsf\{a\}^\{g\}\_\{i\}\|\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\)\}\}z\_\{ji,l\}\\mathbf\{m\}^\{g\}\_\{ji,l\}\\biggr\],\(24\)๐’ดi,l\\displaystyle\\mathcal\{Y\}\_\{i,l\}=\[๐ฑig,โ„‹i,lโˆ’1gโŸfrom the ligand,โ„‹i,lโˆ’1c,โ„‹i,lโˆ’1a,โŸfrom the pocketโˆ‘๐–บjgโˆˆ๐\(๐–บig\|๐’Ÿ\)zjโ€‹i,lMjโ€‹i,lg\],\\displaystyle=\\biggr\[\\underbrace\{\\mathbf\{x\}^\{g\}\_\{i\},\\mathcal\{H\}^\{g\}\_\{i,l\-1\}\}\_\{\\mathclap\{\\text\{from the ligand\}\}\},~\\underbrace\{\\mathcal\{H\}^\{c\}\_\{i,l\-1\},\\mathcal\{H\}^\{a\}\_\{i,l\-1\},\}\_\{\\mathclap\{\\text\{from the pocket\}\}\}\\sum\_\{\\mathsf\{a\}^\{g\}\_\{j\}\\in\{\\mbox\{$\\mathop\{\\mathbf\{N\}\}\\limits$\}\(\\mathsf\{a\}^\{g\}\_\{i\}\|\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\)\}\}z\_\{ji,l\}M^\{g\}\_\{ji,l\}\\biggr\],wherecrepresents embeddings that model pocket\-ligand atomic interactions, andarepresents residue\-level interaction embeddings\.๐ฌi,lโˆ’1c\\mathbf\{s\}^\{c\}\_\{i,l\-1\}andโ„‹i,lโˆ’1c\\mathcal\{H\}^\{c\}\_\{i,l\-1\}denote the scalar and vector embeddings that encode how ligand atom๐–บig\\mathsf\{a\}^\{g\}\_\{i\}interacts with neighboring pocket atoms from the\(lโˆ’1\)\(l\-1\)\-th layer of the GNN \(detailed in Equation[29](https://arxiv.org/html/2607.12349#Sx6.E29)\);๐ฌi,lโˆ’1a\\mathbf\{s\}^\{a\}\_\{i,l\-1\}andโ„‹i,lโˆ’1a\\mathcal\{H\}^\{a\}\_\{i,l\-1\}denote the scalar and vector embeddings that capture structures from surrounding pocket residues for ligand atom๐–บig\\mathsf\{a\}^\{g\}\_\{i\}\(detailed in Equation[30](https://arxiv.org/html/2607.12349#Sx6.E30)\);๐ฆjโ€‹i,lg\\mathbf\{m\}^\{g\}\_\{ji,l\}andMjโ€‹i,lgM^\{g\}\_\{ji,l\}represent the scalar and vector message embeddings in the\(lโˆ’1\)\(l\-1\)\-th layer of the GNN to propagate information from๐–บig\\mathsf\{a\}^\{g\}\_\{i\}โ€™s neighboring atom๐–บjg\\mathsf\{a\}^\{g\}\_\{j\}in the ligand๐’Ÿ\\mathop\{\\mathcal\{D\}\}\\limits\(i\.e\.,๐\(๐–บig\|๐’Ÿ\)\\mbox\{$\\mathop\{\\mathbf\{N\}\}\\limits$\}\(\\mathsf\{a\}^\{g\}\_\{i\}\|\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\)\) to๐–บig\\mathsf\{a\}^\{g\}\_\{i\};zjโ€‹i,lz\_\{ji,l\}is the attention weights modeled in a similar way as in Equation[5](https://arxiv.org/html/2607.12349#Sx6.E5)\. The convolution in Equation[23](https://arxiv.org/html/2607.12349#Sx6.E23)combines information from ligand atom neighbors and pocket atom/residue neighbors, allowing each ligand atom to refine its representation with structural and chemical information from both the ligand and the pocket\.๐ฆjโ€‹i,lg\\mathbf\{m\}^\{g\}\_\{ji,l\}andMjโ€‹i,lgM^\{g\}\_\{ji,l\}are updated as follows:

\(๐ฆjโ€‹i,lg,Mjโ€‹i,lg\)=GVPโ€‹\(๐ฆ^jโ€‹i,lg,M^jโ€‹i,lg\),where\(\\mathbf\{m\}^\{g\}\_\{ji,l\},M^\{g\}\_\{ji,l\}\)=\\text\{GVP\}\(\{\{\\hat\{\\mathbf\{m\}\}^\{g\}\_\{ji,l\}\}\},\{\{\\hat\{M\}^\{g\}\_\{ji,l\}\}\}\),\\text\{ where \}\(25\)๐ฆ^jโ€‹i,lg\\displaystyle\{\\hat\{\\mathbf\{m\}\}^\{g\}\_\{ji,l\}\}=\[๐ฆjโ€‹i,lโˆ’1g,dโ€‹\(๐–บjg,๐–บig\),๐žjโ€‹i,lโˆ’1\],\\displaystyle=\[\{\\mathbf\{m\}\}^\{g\}\_\{ji,l\-1\},\{d\(\\mathsf\{a\}^\{g\}\_\{j\},\\mathsf\{a\}^\{g\}\_\{i\}\)\},\_\{ji,l\-1\}\],\(26\)M^jโ€‹i,lg\\displaystyle\\quad\{\\hat\{M\}^\{g\}\_\{ji,l\}\}=\[Mjโ€‹i,lโˆ’1g,๐ฑjgโˆ’๐ฑig\],\\displaystyle=\[M^\{g\}\_\{ji,l\-1\},\{\\mathbf\{x\}^\{g\}\_\{j\}\-\\mathbf\{x\}^\{g\}\_\{i\}\}\],where๐žjโ€‹i,lโˆ’1\\mbox\{$\\mathop\{\\mathbf\{e\}\}\\limits$\}\_\{ji,l\-1\}is the embedding of the bond type between๐–บig\\mathsf\{a\}^\{g\}\_\{i\}and๐–บjg\\mathsf\{a\}^\{g\}\_\{j\}\(detailed in Equation[27](https://arxiv.org/html/2607.12349#Sx6.E27)\), anddโ€‹\(๐–บjg,๐–บig\)d\(\\mathsf\{a\}^\{g\}\_\{j\},\\mathsf\{a\}^\{g\}\_\{i\}\)is the distance between๐–บig\\mathsf\{a\}^\{g\}\_\{i\}and๐–บjg\\mathsf\{a\}^\{g\}\_\{j\}, and\(๐ฑjgโˆ’๐ฑig\)\(\\mathbf\{x\}^\{g\}\_\{j\}\-\\mathbf\{x\}^\{g\}\_\{i\}\)represents the displacement vector from๐–บig\\mathsf\{a\}^\{g\}\_\{i\}to๐–บjg\\mathsf\{a\}^\{g\}\_\{j\}\.๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitsutilizes bond type embeddings๐žjโ€‹i,l\\mbox\{$\\mathop\{\\mathbf\{e\}\}\\limits$\}\_\{ji,l\}to facilitate its understanding of relations among atoms\. Particularly,๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitsgenerates these bond type embeddings as follows:

๐žjโ€‹i,l=\{MLPโ€‹\(\[๐ฌi,lg\+๐ฌj,lg,absโ€‹\(๐ฌi,lgโˆ’๐ฌj,lg\),dโ€‹\(๐–บjg,๐–บig\)\]\),ifโ€‹l=0,MLPโ€‹\(\[๐ฌi,lg\+๐ฌj,lg,absโ€‹\(๐ฌi,lgโˆ’๐ฌj,lg\),โ€–โ„‹i,lgโ€–2\+โ€–โ„‹j,lgโ€–2,absโ€‹\(โ€–โ„‹i,lgโ€–2โˆ’โ€–โ„‹j,lgโ€–2\)\]\),otherwise\.\\mbox\{$\\mathop\{\\mathbf\{e\}\}\\limits$\}\_\{ji,l\}=\\begin\{cases\}\\text\{MLP\}\(\[\\mathbf\{s\}^\{g\}\_\{i,l\}\+\\mathbf\{s\}^\{g\}\_\{j,l\},\\text\{abs\}\(\\mathbf\{s\}^\{g\}\_\{i,l\}\-\\mathbf\{s\}^\{g\}\_\{j,l\}\),d\(\\mathsf\{a\}^\{g\}\_\{j\},\\mathsf\{a\}^\{g\}\_\{i\}\)\]\),&\\text\{if\}\\;l=0,\\\\ \\text\{MLP\}\(\[\\mathbf\{s\}^\{g\}\_\{i,l\}\+\\mathbf\{s\}^\{g\}\_\{j,l\},\\text\{abs\}\(\\mathbf\{s\}^\{g\}\_\{i,l\}\-\\mathbf\{s\}^\{g\}\_\{j,l\}\),\\\|\\mathcal\{H\}^\{g\}\_\{i,l\}\\\|\_\{2\}\+\\\|\\mathcal\{H\}^\{g\}\_\{j,l\}\\\|\_\{2\},\\text\{abs\}\(\\\|\\mathcal\{H\}^\{g\}\_\{i,l\}\\\|\_\{2\}\-\\\|\\mathcal\{H\}^\{g\}\_\{j,l\}\\\|\_\{2\}\)\]\),&\\text\{ otherwise\}\.\\end\{cases\}\(27\)where abs\(โ‹…\\cdot\) represents the absolute difference\. Intuitively, the bond type embeddings are constructed to be agnostic to the order of atoms๐–บig\\mathsf\{a\}^\{g\}\_\{i\}and๐–บjg\\mathsf\{a\}^\{g\}\_\{j\}by using two invariant operations: the sum and the absolute difference operation\. The sum combines features from neighboring atoms๐–บig\\mathsf\{a\}^\{g\}\_\{i\}and๐–บjg\\mathsf\{a\}^\{g\}\_\{j\}, and the absolute difference estimates the distance between๐–บig\\mathsf\{a\}^\{g\}\_\{i\}and๐–บjg\\mathsf\{a\}^\{g\}\_\{j\}on latent space, offering a comprehensive representation of the bond\. AfterLLlayers,๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitsestimates the atom position and type for each ligand atom๐–บig\\mathsf\{a\}^\{g\}\_\{i\}using the learned atom embeddings as follows:

๐ฑ~ig=๐ฑig\+โ„‹i,Lg,๐ฏ~ig=softmaxโ€‹\(MLPโ€‹\(๐ฌi,Lg\)\),\{\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{i\}=\\mathbf\{x\}^\{g\}\_\{i\}\+\\mathcal\{H\}^\{g\}\_\{i,L\}\},\\quad\\tilde\{\\mathbf\{v\}\}^\{g\}\_\{i\}=\\text\{softmax\}\(\\text\{MLP\}\(\\mathbf\{s\}^\{g\}\_\{i,L\}\)\),\(28\)where๐ฑ~ig\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{i\}and๐ฑig\\mathbf\{x\}^\{g\}\_\{i\}are the predicted noise\-free atom positions and input noisy atom positions of atom๐–บig\\mathsf\{a\}^\{g\}\_\{i\}, respectively;โ„‹i,Lgโˆˆโ„1ร—3\\mathcal\{H\}^\{g\}\_\{i,L\}\\in\\mathbb\{R\}^\{1\\times 3\}is the predicted noise for๐–บig\\mathsf\{a\}^\{g\}\_\{i\}, and๐ฌi,Lg\\mathbf\{s\}^\{g\}\_\{i,L\}is used for atom type prediction of๐–บig\\mathsf\{a\}^\{g\}\_\{i\}\.๐ฏ~ig\\tilde\{\\mathbf\{v\}\}^\{g\}\_\{i\}represents the predicted clean categorical distribution of atom features\.

#### Pocket\-ligand Interaction Learning

This section describes how to learn pocket\-ligand interaction embeddings \(embeddings with superscriptcorrin Equation[24](https://arxiv.org/html/2607.12349#Sx6.E24)\) by leveraging pocket embeddings \(embeddings with superscriptp\) encoded by๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits\. For each ligand atom๐–บig\\mathsf\{a\}^\{g\}\_\{i\}, at eachll\-th layer of the GNN,๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitslearns two invariant embeddings๐ฌi,lc\\mathbf\{s\}^\{c\}\_\{i,l\}and๐ฌi,la\\mathbf\{s\}^\{a\}\_\{i,l\}, and two equivariant embeddingโ„‹i,lc\\mathcal\{H\}^\{c\}\_\{i,l\}andโ„‹i,la\\mathcal\{H\}^\{a\}\_\{i,l\}, encoding how๐–บig\\mathsf\{a\}^\{g\}\_\{i\}interacts with neighboring pocket atoms and residues\. These embeddings are learned in a recurrent manner as follows:

๐ฌi,lc=MLPโ€‹\(\[๐ฌ~i,lc,๐ฌi,lโˆ’1c\]\),โ„‹i,lc=VN\-MLPโ€‹\(\[โ„‹~i,lc,โ„‹i,lโˆ’1c\]\),\\mathbf\{s\}^\{c\}\_\{i,l\}=\\text\{MLP\}\(\[\\tilde\{\\mathbf\{s\}\}^\{c\}\_\{i,l\},\\mathbf\{s\}^\{c\}\_\{i,l\-1\}\]\),~~\\mathcal\{H\}^\{c\}\_\{i,l\}=\\text\{VN\-MLP\}\(\[\\tilde\{\\mathcal\{H\}\}^\{c\}\_\{i,l\},\\mathcal\{H\}^\{c\}\_\{i,l\-1\}\]\),\(29\)๐ฌi,la=MLPโ€‹\(\[๐ฌ~i,la,๐ฌi,lโˆ’1a\]\),โ„‹i,la=VN\-MLPโ€‹\(\[โ„‹~i,la,โ„‹i,lโˆ’1a\]\)\.\\mathbf\{s\}^\{a\}\_\{i,l\}=\\text\{MLP\}\(\[\\tilde\{\\mathbf\{s\}\}^\{a\}\_\{i,l\},\\mathbf\{s\}^\{a\}\_\{i,l\-1\}\]\),~~\\mathcal\{H\}^\{a\}\_\{i,l\}=\\text\{VN\-MLP\}\(\[\\tilde\{\\mathcal\{H\}\}^\{a\}\_\{i,l\},\\mathcal\{H\}^\{a\}\_\{i,l\-1\}\]\)\.\(30\)That is,๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitsupdates the atom\-level\(๐ฌic,โ„‹ic\)\(\\mathbf\{s\}^\{c\}\_\{i\},\\mathcal\{H\}^\{c\}\_\{i\}\)and residue\-level\(๐ฌia,โ„‹ia\)\(\\mathbf\{s\}^\{a\}\_\{i\},\\mathcal\{H\}^\{a\}\_\{i\}\)pocket\-ligand interaction embeddings recurrently across layers\. We model๐ฌ~i,lc\\tilde\{\\mathbf\{s\}\}^\{c\}\_\{i,l\}andโ„‹~i,lc\\tilde\{\\mathcal\{H\}\}^\{c\}\_\{i,l\}as follows:

\(๐ฌ~i,lc,โ„‹~i,lc\)=GVPโ€‹\(๐›i,lc,โ„i,lc\),where\(\\tilde\{\\mathbf\{s\}\}^\{c\}\_\{i,l\},\\tilde\{\\mathcal\{H\}\}^\{c\}\_\{i,l\}\)=\\text\{GVP\}\(\\mathbf\{b\}\_\{i,l\}^\{c\},\\mathcal\{I\}\_\{i,l\}^\{c\}\),\\text\{ where \}\(31\)๐›i,lc\\displaystyle\\mathbf\{b\}\_\{i,l\}^\{c\}=โˆ‘jโˆˆ๐\(๐–บig\|๐’ซ\)nwiโ€‹j,lcโ€‹MLPโ€‹\(\[dโ€‹\(๐–บig,๐–บjp\),๐ฌi,lg,๐ฌjp\]\),\\displaystyle=\\sum\_\{j\\in\{\\mbox\{$\\mathop\{\\mathbf\{N\}\}\\limits$\}\(\\mathsf\{a\}^\{g\}\_\{i\}\|\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\)\}\}^\{n\}w\_\{ij,l\}^\{c\}\\text\{MLP\}\(\[d\(\\mathsf\{a\}^\{g\}\_\{i\},\\mathsf\{a\}^\{p\}\_\{j\}\),\\mathbf\{s\}^\{g\}\_\{i,l\},\\mathbf\{s\}^\{p\}\_\{j\}\]\),\(32\)โ„i,lc\\displaystyle\\mathcal\{I\}\_\{i,l\}^\{c\}=โˆ‘๐–บjpโˆˆ๐\(๐–บig\|๐’ซ\)nwiโ€‹j,lcโ€‹VN\-MLPโ€‹\(\[๐ฑjpโˆ’๐ฑig,โ„‹i,lg,โ„‹jp\]\),\\displaystyle=\\sum\_\{\\mathsf\{a\}^\{p\}\_\{j\}\\in\{\\mbox\{$\\mathop\{\\mathbf\{N\}\}\\limits$\}\(\\mathsf\{a\}^\{g\}\_\{i\}\|\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\)\}\}^\{n\}w\_\{ij,l\}^\{c\}\\text\{VN\-MLP\}\(\[\\mathbf\{x\}^\{p\}\_\{j\}\-\\mathbf\{x\}^\{g\}\_\{i\},\\mathcal\{H\}^\{g\}\_\{i,l\},\\mathcal\{H\}^\{p\}\_\{j\}\]\),where๐’ซ\\mathop\{\\mathcal\{P\}\}\\limitsis the target pocket of ligand๐’Ÿ\\mathop\{\\mathcal\{D\}\}\\limits, and๐\(๐–บig\|๐’ซ\)\{\\mbox\{$\\mathop\{\\mathbf\{N\}\}\\limits$\}\(\\mathsf\{a\}^\{g\}\_\{i\}\|\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\)\}denotes thenn\-nearest pocket atoms surrounding the ligand atom๐–บig\\mathsf\{a\}^\{g\}\_\{i\};๐ฌjp\\mathbf\{s\}^\{p\}\_\{j\}andโ„‹jp\\mathcal\{H\}^\{p\}\_\{j\}are pretrained pocket atom embeddings \(Equation[3](https://arxiv.org/html/2607.12349#Sx6.E3)\);๐ฌi,lโˆ’1g\\mathbf\{s\}^\{g\}\_\{i,l\-1\}andโ„‹i,lโˆ’1g\\mathcal\{H\}^\{g\}\_\{i,l\-1\}are ligand atom embeddings as presented in Equation[23](https://arxiv.org/html/2607.12349#Sx6.E23); anddโ€‹\(๐–บig,๐–บjp\)d\(\\mathsf\{a\}^\{g\}\_\{i\},\\mathsf\{a\}^\{p\}\_\{j\}\)is the Euclidean distance between๐–บig\\mathsf\{a\}^\{g\}\_\{i\}and๐–บjp\\mathsf\{a\}^\{p\}\_\{j\}\. The attention weights are determined based on the ligand embeddings \(Equation[23](https://arxiv.org/html/2607.12349#Sx6.E23)\) and pocket embeddings provided by๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limitsas follows:

wiโ€‹j,lc\\displaystyle w\_\{ij,l\}^\{c\}=softmaxโ€‹\(MLPโ€‹\(\[dโ€‹\(๐–บig,๐–บjp\),๐ฏiโ€‹j,lc,๐ฌi,lg,๐ฌjp\]\)\),\\displaystyle=\\text\{softmax\}\(\\text\{MLP\}\(\[d\(\\mathsf\{a\}^\{g\}\_\{i\},\\mathsf\{a\}^\{p\}\_\{j\}\),\\mathbf\{v\}^\{c\}\_\{ij,l\},\\mathbf\{s\}^\{g\}\_\{i,l\},\\mathbf\{s\}^\{p\}\_\{j\}\]\)\),\(33\)๐ฏiโ€‹j,lc\\displaystyle\\mathbf\{v\}\_\{ij,l\}^\{c\}=โ€–VN\-MLPโ€‹\(\[๐ฑjpโˆ’๐ฑig,โ„‹i,lg,โ„‹jp\]\)โ€–2,\\displaystyle=\\\|\\text\{VN\-MLP\}\(\[\\mathbf\{x\}^\{p\}\_\{j\}\-\\mathbf\{x\}^\{g\}\_\{i\},\\mathcal\{H\}^\{g\}\_\{i,l\},\\mathcal\{H\}^\{p\}\_\{j\}\]\)\\\|\_\{2\},where๐ฏiโ€‹j,l\\mathbf\{v\}\_\{ij,l\}encodes the spatial interactions between๐–บig\\mathsf\{a\}^\{g\}\_\{i\}and๐–บjp\\mathsf\{a\}^\{p\}\_\{j\}\. While\(๐ฌ~i,lc,โ„‹~i,lc\)\(\\tilde\{\\mathbf\{s\}\}^\{c\}\_\{i,l\},\\tilde\{\\mathcal\{H\}\}^\{c\}\_\{i,l\}\)accounts for atomic pocket\-ligand interactions, we also model interactions between ligand atoms and pocket residues\.\(๐ฌ~i,la,โ„‹~i,la\)\(\\tilde\{\\mathbf\{s\}\}^\{a\}\_\{i,l\},\\tilde\{\\mathcal\{H\}\}^\{a\}\_\{i,l\}\)are obtained in a similar manner by aggregating interactions between๐–บig\\mathsf\{a\}^\{g\}\_\{i\}andmm\-nearest pocket residue๐—‹jp\\mathsf\{r\}^\{p\}\_\{j\}\. We summarize the ligand generation procedure of๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsincorporating๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limitsand๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitsin Algorithm[2](https://arxiv.org/html/2607.12349#alg2)\.

Algorithm 2๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsfor ligand generationRequired Input: Ligand๐’Ÿ\\mathop\{\\mathcal\{D\}\}\\limits, pocket๐’ซ\\mathop\{\\mathcal\{P\}\}\\limits

1:

n=sampleNumAtomsโ€‹\(๐’Ÿ,๐’ซ\)n=\\text\{\{sampleNumAtoms\}\}\(\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\},\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\)โŠณ\\trianglerightsample the number of ligand atoms from pocket size

2:

\{๐ฑTg\}nโˆผ๐’ฉโ€‹\(0,๐ˆ\)\\\{\\mathbf\{x\}^\{g\}\_\{T\}\\\}^\{n\}\\sim\\mathcal\{N\}\(0,\\mathbf\{I\}\)โŠณ\\trianglerightinitialize positions of n ligand atoms

3:

\{๐ฏTg\}nโˆผ๐’žโ€‹\(K,1K\)\\\{\\mathbf\{v\}^\{g\}\_\{T\}\\\}^\{n\}\\sim\\mathcal\{C\}\(K,\\frac\{1\}\{K\}\)โŠณ\\trianglerightinitialize types of n ligand atoms

4:

๐ฌp,โ„‹p=๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\(๐ฑp,๐ฏp,๐ฑr,๐ฏr\)\\mathbf\{s\}^\{p\},\\mathcal\{H\}^\{p\}=\\mbox\{$\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits$\}\(\\mathbf\{x\}^\{p\},\\mathbf\{v\}^\{p\},\\mathbf\{x\}^\{r\},\\mathbf\{v\}^\{r\}\)โŠณ\\trianglerightencode pocket into embeddings using๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits

5:for

t=Tt=Tto

11do

6:

\(๐ฑ~0,tg,๐ฑ~0,tg\)=๐—‰๐–ผ๐–ซ๐–ฆ\(๐ฑtg,๐ฏtg,๐ฌp,โ„‹p\)\(\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{0,t\},\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{0,t\}\)=\\mbox\{$\\mathop\{\\mathsf\{pcLG\}\}\\limits$\}\(\\mathbf\{x\}^\{g\}\_\{t\},\\mathbf\{v\}^\{g\}\_\{t\},\\mathbf\{s\}^\{p\},\\mathcal\{H\}^\{p\}\)โŠณ\\trianglerightpredict noise\-free ligand using๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limits

7:

๐ฑtโˆ’1g=qโ€‹\(๐ฑtโˆ’1g\|๐ฑtg,๐ฑ~0,tg\)\\mathbf\{x\}^\{g\}\_\{t\-1\}=q\(\\mathbf\{x\}^\{g\}\_\{t\-1\}\|\\mathbf\{x\}^\{g\}\_\{t\},\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{0,t\}\)โŠณ\\trianglerightsample๐ฑtโˆ’1g\\mathbf\{x\}^\{g\}\_\{t\-1\}using Gaussian posterior \(Equation[16](https://arxiv.org/html/2607.12349#Sx6.E16)\)

8:

๐ฏtโˆ’1g=qโ€‹\(๐ฏtโˆ’1g\|๐ฏtg,๐ฑ~0,tg\)\\mathbf\{v\}^\{g\}\_\{t\-1\}=q\(\\mathbf\{v\}^\{g\}\_\{t\-1\}\|\\mathbf\{v\}^\{g\}\_\{t\},\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{0,t\}\)โŠณ\\trianglerightsample๐ฏtโˆ’1g\\mathbf\{v\}^\{g\}\_\{t\-1\}using categorical posterior \(Equation[17](https://arxiv.org/html/2607.12349#Sx6.E17)\)

9:endfor

10:

๐’Ÿgโ€‹eโ€‹n=\(๐ฑ0g,๐ฏ0g\)\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\_\{gen\}=\(\\mathbf\{x\}^\{g\}\_\{0\},\\mathbf\{v\}^\{g\}\_\{0\}\)
11:return

๐’Ÿgโ€‹eโ€‹n\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\_\{gen\}โŠณ\\trianglerightreturn the generated ligand

### ๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitswith Generation\-time Property\-aware Optimization \(๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\)

๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsis well tailored for pocket\-specific ligand generation\. In this section, we develop a new method that further improves the ligands generated by๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsto exhibit additional desired properties \(e\.g\., ADMET\)\. To this end, we extend๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsby introducing an iterative, generation\-time property\-aware optimization procedure that refines the generated ligands with respect to these desired properties\. In essence, this optimization process โ€œalignsโ€ the model with desired properties beyond binding affinity or drug\-likeness\. We formulate this โ€œalignmentโ€ process as a generation\-time optimization problem, in which the objective is to steer the generated ligands toward exhibiting more favorable properties, such as higher BBBP or lower carcinogenicity\. The proposed optimization scheme, named๐—‰๐–บ๐–ฎ๐–ฏ๐–ณ\\mathop\{\\mathsf\{paOPT\}\}\\limits, operates entirely during the modelโ€™s inference process\. That is, all model parameters of๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsare kept fixed at this stage\. Without requiring any additional model training, the approach is both computationally efficient and readily adaptable to different optimization objectives while preserving๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsโ€™s original generative performance on general molecular metrics\. We refer to the resulting framework that combines๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitswith๐—‰๐–บ๐–ฎ๐–ฏ๐–ณ\\mathop\{\\mathsf\{paOPT\}\}\\limitsas๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\.

#### Alignment Problem Formulation

In๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limits, the stochasticity of ligand generation arises from noise trajectories of atom positions and atom types \(๐ณt๐ฑ\\mathbf\{z\}^\{\\mathbf\{x\}\}\_\{t\}in Equation[16](https://arxiv.org/html/2607.12349#Sx6.E16)and๐ณt๐ฏ\\mathbf\{z\}^\{\\mathbf\{v\}\}\_\{t\}in Equation[17](https://arxiv.org/html/2607.12349#Sx6.E17)\)\. Noise trajectories govern the denoising dynamics throughout the diffusion process, determining how ligand structures evolve over time and ultimately shaping the final generated ligand\.

Therefore, with the diffusion model parameters fixed, the entire generation process can be viewed as a deterministic mapping that takes the noise trajectories as input and outputs the generated ligand\. These noise trajectories constitute the optimization variables in๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\. After generation, the ligand is evaluated by an evaluator functionRโ€‹\(โ‹…\)R\(\\cdot\), which defines the objective function for the optimization problem\. We then employ gradient\-based optimization to iteratively update the noise trajectories, using the gradients of the objective function to steer the generation process towards ligands with more favorable properties\.

##### Noise as Optimization Variables

The generation process of๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limits, described in Algorithm[2](https://arxiv.org/html/2607.12349#alg2), consists ofTTdenoising steps\. It begins with atom positions initialized as pure Gaussian noise,๐ฑi,Tgโˆผ๐’ฉโ€‹\(๐ŸŽ,๐ˆ\)\\mathbf\{x\}^\{g\}\_\{i,T\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)\(here, we define๐ณi,T๐šก=๐ฑi,Tg\\mathbf\{z\}^\{\\mathtt\{x\}\}\_\{i,T\}=\\mathbf\{x\}^\{g\}\_\{i,T\}\), and atom types sampled uniformly at random,๐ฏi,TgโˆผCโ€‹\(K,1K\)\.\\mathbf\{v\}^\{g\}\_\{i,T\}\\sim C\(K,\\frac\{1\}\{K\}\)\.The uniform categorical sampling is implemented via the Gumbel\-max trick\. Specifically, for each atomiiand typekk, in๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits, we draw๐ณi,T,k๐šŸโˆผGumbelโ€‹\(0,1\)\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{i,T,k\}\\sim\\mathrm\{Gumbel\}\(0,1\)fork=1,โ€ฆ,Kk=1,\\dots,K, and set๐ฏi,Tg=argโกmaxkโก\{๐ณi,T,k๐šŸ\}\\mathbf\{v\}^\{g\}\_\{i,T\}=\\arg\\max\_\{k\}\\\{\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{i,T,k\}\\\}\. At each subsequent steptt, the atom positions๐ฑi,tg\\mathbf\{x\}^\{g\}\_\{i,t\}and types๐ฏi,tg\\mathbf\{v\}^\{g\}\_\{i,t\}are updated according to Equations[16](https://arxiv.org/html/2607.12349#Sx6.E16)and[17](https://arxiv.org/html/2607.12349#Sx6.E17), where Gaussian noise๐ณi,tโˆ’1๐šกโˆผ๐’ฉโ€‹\(๐ŸŽ,๐ˆ\)\\mathbf\{z\}^\{\\mathtt\{x\}\}\_\{i,t\-1\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)and Gumbel noises๐ณi,tโˆ’1,k๐šŸโˆผGumbel\(0,1\)\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{i,t\-1,k\}\\sim\\text\{Gumbel\(0,1\)\}fork=1,โ€ฆ,Kk=1,\\dots,Kare injected into the denoising updates to produce the predicted๐ฑi,tโˆ’1g\\mathbf\{x\}^\{g\}\_\{i,t\-1\}and๐ฏi,tโˆ’1g\\mathbf\{v\}^\{g\}\_\{i,t\-1\}\. For compactness, we write\[๐ณi,t,1๐šŸ,โ€ฆ,๐ณi,t,k๐šŸ\]\[\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{i,t,1\},\\dots,\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{i,t,k\}\]as๐ณi,t๐šŸ\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{i,t\}\. We collect these noises into two noise trajectories: Gaussian noise trajectory\{๐ณi,T๐šก,๐ณi,Tโˆ’1๐šก,โ€ฆ,๐ณi,0๐šก\}\\\{\\mathbf\{z\}^\{\\mathtt\{x\}\}\_\{i,T\},\\mathbf\{z\}^\{\\mathtt\{x\}\}\_\{i,T\-1\},\\dots,\\mathbf\{z\}^\{\\mathtt\{x\}\}\_\{i,0\}\\\}and Gumbel noise trajectory\{๐ณi,T๐šŸ,๐ณi,T๐šŸ,โ€ฆ,๐ณi,0๐šŸ\}\\\{\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{i,T\},\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{i,T\},\\dots,\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{i,0\}\\\}\. By treating the noise trajectories as optimization variables, we aim to iteratively refine them to guide the generative process towards chemical space with improved target properties\.

##### Objective Function

LetD0D\_\{0\}represent the generated ligand at the final timestep\. To evaluate its quality with respect to a desired ADMET property, we employ a property\-specific evaluator function\. We denote the evaluator asRโ€‹\(โ‹…\)R\(\\cdot\), which serves as the objective function in our formulation\. The choice of evaluator is flexible and can be tailored to specific target properties\. For example, if the goal is to improve the blood\-brain barrier permeability of the generated ligandD0D\_\{0\}, thenRโ€‹\(D0\)R\(D\_\{0\}\)can be defined as the predicted probability \(given by an external property classifier\) that ligandD0D\_\{0\}is blood\-brain barrier permeable\.

##### Noise Optimization

For notational simplicity, we denote the entire Gaussian noise trajectory\{๐ณi,T๐šก,๐ณi,Tโˆ’1๐šก,โ€ฆ,๐ณi,0๐šก\}\\\{\\mathbf\{z\}^\{\\mathtt\{x\}\}\_\{i,T\},\\mathbf\{z\}^\{\\mathtt\{x\}\}\_\{i,T\-1\},\\dots,\\mathbf\{z\}^\{\\mathtt\{x\}\}\_\{i,0\}\\\}asฯ„๐ฑ\\tau^\{\\mathbf\{x\}\}and the entire Gumbel noise trajectory\{๐ณi,T๐šŸ,๐ณi,Tโˆ’1๐šŸ,โ€ฆ,๐ณi,0๐šŸ\}\\\{\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{i,T\},\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{i,T\-1\},\\dots,\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{i,0\}\\\}asฯ„๐ฏ\\tau^\{\\mathbf\{v\}\}\. With๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitsfixed,ฯ„๐ฑ\\tau^\{\\mathbf\{x\}\}andฯ„๐ฏ\\tau^\{\\mathbf\{v\}\}fully determine the generated ligand๐’Ÿ0\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\_\{0\}\. Therefore, we write๐’Ÿ0=Mโ€‹\(ฯ„๐ฑ,ฯ„๐ฏ\)\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\_\{0\}=M\(\\tau^\{\\mathbf\{x\}\},\\tau^\{\\mathbf\{v\}\}\), whereMMdenotes the deterministic mapping from noise trajectories to the generated ligand\. Our goal is to adjust the noise trajectoriesฯ„๐ฑ\\tau^\{\\mathbf\{x\}\}andฯ„๐ฏ\\tau^\{\\mathbf\{v\}\}so that the generated ligand has a more favorable property scoreRR\. The generation\-time noise optimization problem is thus formulated as:

maxฯ„๐ฑ,ฯ„๐ฏโกRโ€‹\(Mโ€‹\(ฯ„๐ฑ,ฯ„๐ฏ\)\)\\max\_\{\\tau^\{\\mathbf\{x\}\},\\tau^\{\\mathbf\{v\}\}\}R\(M\(\\tau^\{\\mathbf\{x\}\},\\tau^\{\\mathbf\{v\}\}\)\)\(34\)

#### ๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsOptimization

We adopt a straightforward gradient\-based method, as presented in Algorithm[3](https://arxiv.org/html/2607.12349#alg3), to solve the noise optimization problem \(Equation[34](https://arxiv.org/html/2607.12349#Sx6.E34)\)\. We use the noise trajectories produced by๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsas a starting point to be optimized\. The initial noise trajectories are iteratively refined through gradient\-based updates\. Specifically, at each iteration, we evaluate the objectiveRโ€‹\(Mโ€‹\(ฯ„๐ฑ,ฯ„๐ฏ\)\)R\(M\(\\tau^\{\\mathbf\{x\}\},\\tau^\{\\mathbf\{v\}\}\)\), calculate the gradient ofRRwith respect toฯ„๐ฑ,ฯ„๐ฏ\\tau^\{\\mathbf\{x\}\},\\tau^\{\\mathbf\{v\}\}, and update the noise accordingly via gradient ascent\. This approach is applicable when the gradient of the objective functionRRis available\. However, in practice, such gradients are often inaccessible, since many ADMET evaluators operate as black\-box predictors that provide only scalar property scores\. To overcome this limitation, we employ zero*th*\-order gradient estimation, enabling gradient\-based updates without explicit analytic gradients\. Algorithm[5](https://arxiv.org/html/2607.12349#alg5)presents the๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsalgorithm\.

Algorithm 3๐–ญ๐—ˆ๐—‚๐—Œ๐–พ๐–ฎ๐—‰๐—\\mathop\{\\mathsf\{NoiseOpt\}\}\\limitsfor noise trajectory optimizationRequired Input: Initial noise trajectoriesฯ„0๐ฑ,ฯ„0๐ฏ\\tau^\{\\mathbf\{x\}\}\_\{0\},\\tau^\{\\mathbf\{v\}\}\_\{0\}, iteration numberNN, step sizeฮฑ\\alpha

1:

rbestโ†Rโ€‹\(Mโ€‹\(ฯ„0๐ฑ,ฯ„0๐ฏ\)\),nbestโ†0r\_\{\\mathrm\{best\}\}\\leftarrow R\(M\(\\tau^\{\\mathbf\{x\}\}\_\{0\},\\tau^\{\\mathbf\{v\}\}\_\{0\}\)\),n\_\{\\mathrm\{best\}\}\\leftarrow 0โŠณ\\trianglerightevaluate objective at initialization

2:for

n=0n=0to

Nโˆ’1N\-1do

3:

rโ†Rโ€‹\(Mโ€‹\(ฯ„n๐ฑ,ฯ„n๐ฏ\)\)r\\leftarrow R\(M\(\\tau^\{\\mathbf\{x\}\}\_\{n\},\\tau^\{\\mathbf\{v\}\}\_\{n\}\)\)โŠณ\\trianglerightevaluate current objective

4:if

r\>rbestr\>r\_\{\\mathrm\{best\}\}then

5:

rbestโ†r,nbestโ†nr\_\{\\mathrm\{best\}\}\\leftarrow r,n\_\{\\mathrm\{best\}\}\\leftarrow nโŠณ\\trianglerightupdate best step

6:endif

7:

โˆ‡^โ€‹Rโ€‹\(Mโ€‹\(ฯ„n๐ฑ,ฯ„n๐ฏ\)\)=๐–น๐–ฎ๐–ฆ๐—‹๐–บ๐–ฝ\(ฯ„n๐ฑ,ฯ„n๐ฏ\)\\hat\{\\nabla\}R\(M\(\\tau^\{\\mathbf\{x\}\}\_\{n\},\\tau^\{\\mathbf\{v\}\}\_\{n\}\)\)=\\mbox\{$\\mathop\{\\mathsf\{ZOGrad\}\}\\limits$\}\(\\tau^\{\\mathbf\{x\}\}\_\{n\},\\tau^\{\\mathbf\{v\}\}\_\{n\}\)โŠณ\\trianglerightestimate zeroth\-order gradients via Algorithm[4](https://arxiv.org/html/2607.12349#alg4)

8:

ฯ„n\+1๐ฑ=ฯ„n๐ฑ\+ฮฑโ€‹โˆ‡^ฯ„๐ฑโ€‹Rโ€‹\(Mโ€‹\(ฯ„n๐ฑ,ฯ„n๐ฏ\)\)\\tau^\{\\mathbf\{x\}\}\_\{n\+1\}=\\tau^\{\\mathbf\{x\}\}\_\{n\}\+\\alpha\\hat\{\\nabla\}\_\{\\tau^\{\\mathbf\{x\}\}\}R\(M\(\\tau^\{\\mathbf\{x\}\}\_\{n\},\\tau^\{\\mathbf\{v\}\}\_\{n\}\)\)โŠณ\\trianglerightupdate Gaussian noise trajectory

9:

ฯ„n\+1๐ฏ=ฯ„n๐ฏ\+ฮฑโ€‹โˆ‡^ฯ„๐ฏโ€‹Rโ€‹\(Mโ€‹\(ฯ„n๐ฑ,ฯ„n๐ฏ\)\)\\tau^\{\\mathbf\{v\}\}\_\{n\+1\}=\\tau^\{\\mathbf\{v\}\}\_\{n\}\+\\alpha\\hat\{\\nabla\}\_\{\\tau^\{\\mathbf\{v\}\}\}R\(M\(\\tau^\{\\mathbf\{x\}\}\_\{n\},\\tau^\{\\mathbf\{v\}\}\_\{n\}\)\)โŠณ\\trianglerightupdate Gumbel noise trajectory

10:endfor

11:

๐’Ÿbestโ†Mโ€‹\(ฯ„nbest๐ฑ,ฯ„nbest๐ฏ\)\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\_\{\\mathrm\{best\}\}\\leftarrow M\(\\tau^\{\\mathbf\{x\}\}\_\{n\_\{\\mathrm\{best\}\}\},\\tau^\{\\mathbf\{v\}\}\_\{n\_\{\\mathrm\{best\}\}\}\)
12:return

๐’Ÿbest\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\_\{\\mathrm\{best\}\}โŠณ\\trianglerightreturn the best ligand found among all iterations

##### Zero*th*\-order Gradient Estimation

Algorithm[4](https://arxiv.org/html/2607.12349#alg4)outlines the zero*th*\-order gradient estimation procedure used in๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\. The key idea is to estimate the gradient of the objective function by observing how small random perturbations to the input variables change the output score\. At each iteration, we sample random perturbation directions for both the Gaussian and Gumbel noise trajectories,u๐ฑโˆผ๐’ฉโ€‹\(0,๐ˆ๐ฑ\)u^\{\\mathbf\{x\}\}\\sim\\mathcal\{N\}\(0,\\mathbf\{I\}^\{\\mathbf\{x\}\}\)andu๐ฏโˆผ๐’ฉโ€‹\(0,๐ˆ๐ฏ\)u^\{\\mathbf\{v\}\}\\sim\\mathcal\{N\}\(0,\\mathbf\{I\}^\{\\mathbf\{v\}\}\), and generate perturbed versions of the trajectories:\(ฯ„๐ฑ\+ฮผโ€‹u๐ฑ,ฯ„๐ฏ\+ฮผโ€‹u๐ฏ\)\(\\tau^\{\\mathbf\{x\}\}\+\\mu u^\{\\mathbf\{x\}\},\\tau^\{\\mathbf\{v\}\}\+\\mu u^\{\\mathbf\{v\}\}\)and\(ฯ„๐ฑโˆ’ฮผโ€‹u๐ฑ,ฯ„๐ฏโˆ’ฮผโ€‹u๐ฏ\)\(\\tau^\{\\mathbf\{x\}\}\-\\mu u^\{\\mathbf\{x\}\},\\tau^\{\\mathbf\{v\}\}\-\\mu u^\{\\mathbf\{v\}\}\), whereฮผ\\muis the perturbation magnitude\. The objective function is then evaluated at these perturbed points, and the gradient is estimated from the scaled difference between the corresponding objective values, providing a stochastic estimation of the true gradient direction\. By averaging the estimated directional derivatives over multiple perturbations, we obtain unbiased gradient estimates with respect to bothฯ„๐ฑ\\tau^\{\\mathbf\{x\}\}andฯ„๐ฏ\\tau^\{\\mathbf\{v\}\}\. This enables effective gradient\-based optimization even when the evaluator is a black\-box predictor that provides only function values without analytic gradients\.

Algorithm 4๐–น๐–ฎ๐–ฆ๐—‹๐–บ๐–ฝ\\mathop\{\\mathsf\{ZOGrad\}\}\\limitsfor zeroth\-order gradient estimationRequired Input: Current trajectories\(ฯ„๐ฑ,ฯ„๐ฏ\)\(\\tau^\{\\mathbf\{x\}\},\\tau^\{\\mathbf\{v\}\}\), smoothingฮผ\>0\\mu\>0, number of random perturbationsHH

1:Initialize accumulators:

g๐ฑโ†๐ŸŽg^\{\\mathbf\{x\}\}\\\!\\leftarrow\\mathbf\{0\},

g๐ฏโ†๐ŸŽg^\{\\mathbf\{v\}\}\\\!\\leftarrow\\mathbf\{0\}
2:for

h=1h=1to

HHdo

3:

uh๐ฑโˆผ๐’ฉโ€‹\(0,๐ˆ๐ฑ\)u^\{\\mathbf\{x\}\}\_\{h\}\\sim\\mathcal\{N\}\(0,\\mathbf\{I^\{\\mathbf\{x\}\}\}\),

uh๐ฏโˆผ๐’ฉโ€‹\(0,๐ˆ๐ฏ\)u^\{\\mathbf\{v\}\}\_\{h\}\\sim\\mathcal\{N\}\(0,\\mathbf\{I^\{\\mathbf\{v\}\}\}\)โŠณ\\trianglerightsample random perturbations

4:

rh\+โ†Rโ€‹\(Mโ€‹\(ฯ„๐ฑ\+ฮผโ€‹uh๐ฑ,ฯ„๐ฏ\+ฮผโ€‹uh๐ฏ\)\)r^\{\+\}\_\{h\}\\leftarrow R\\\!\\left\(M\\\!\\big\(\\tau^\{\\mathbf\{x\}\}\+\\mu u^\{\\mathbf\{x\}\}\_\{h\},\\;\\tau^\{\\mathbf\{v\}\}\+\\mu u^\{\\mathbf\{v\}\}\_\{h\}\\big\)\\right\)โŠณ\\trianglerightpositive perturbations

5:

rhโˆ’โ†Rโ€‹\(Mโ€‹\(ฯ„๐ฑโˆ’ฮผโ€‹uh๐ฑ,ฯ„๐ฏโˆ’ฮผโ€‹uh๐ฏ\)\)r^\{\-\}\_\{h\}\\leftarrow R\\\!\\left\(M\\\!\\big\(\\tau^\{\\mathbf\{x\}\}\-\\mu u^\{\\mathbf\{x\}\}\_\{h\},\\;\\tau^\{\\mathbf\{v\}\}\-\\mu u^\{\\mathbf\{v\}\}\_\{h\}\\big\)\\right\)โŠณ\\trianglerightnegative perturbations

6:

ฮ”โ€‹rhโ†rh\+โˆ’rhโˆ’2โ€‹ฮผ\\Delta r\_\{h\}\\leftarrow\\dfrac\{r^\{\+\}\_\{h\}\-r^\{\-\}\_\{h\}\}\{2\\mu\}โŠณ\\trianglerightcalculate directional difference

7:

g๐ฑโ†g๐ฑ\+ฮ”โ€‹rhโ€‹uh๐ฑ,g๐ฏโ†g๐ฏ\+ฮ”โ€‹rhโ€‹uh๐ฏg^\{\\mathbf\{x\}\}\\leftarrow g^\{\\mathbf\{x\}\}\+\\Delta r\_\{h\}\\,u^\{\\mathbf\{x\}\}\_\{h\},\\qquad g^\{\\mathbf\{v\}\}\\leftarrow g^\{\\mathbf\{v\}\}\+\\Delta r\_\{h\}\\,u^\{\\mathbf\{v\}\}\_\{h\}โŠณ\\trianglerightaccumulate estimators

8:endfor

9:

โˆ‡^ฯ„๐ฑโ€‹Rโ€‹\(Mโ€‹\(ฯ„๐ฑ,ฯ„๐ฏ\)\)โ†1Hโ€‹g๐ฑ\\widehat\{\\nabla\}\_\{\\tau^\{\\mathbf\{x\}\}\}R\(M\(\\tau^\{\\mathbf\{x\}\},\\tau^\{\\mathbf\{v\}\}\)\)\\leftarrow\\dfrac\{1\}\{H\}g^\{\\mathbf\{x\}\},

โˆ‡^ฯ„๐ฏโ€‹Rโ€‹\(Mโ€‹\(ฯ„๐ฑ,ฯ„๐ฏ\)\)โ†1Hโ€‹g๐ฏ\\widehat\{\\nabla\}\_\{\\tau^\{\\mathbf\{v\}\}\}R\(M\(\\tau^\{\\mathbf\{x\}\},\\tau^\{\\mathbf\{v\}\}\)\)\\leftarrow\\dfrac\{1\}\{H\}g^\{\\mathbf\{v\}\}
10:Return

โˆ‡^ฯ„๐ฑโ€‹Rโ€‹\(Mโ€‹\(ฯ„๐ฑ,ฯ„๐ฏ\)\)\\widehat\{\\nabla\}\_\{\\tau^\{\\mathbf\{x\}\}\}R\(M\(\\tau^\{\\mathbf\{x\}\},\\tau^\{\\mathbf\{v\}\}\)\),

โˆ‡^ฯ„๐ฏโ€‹Rโ€‹\(Mโ€‹\(ฯ„๐ฑ,ฯ„๐ฏ\)\)\\widehat\{\\nabla\}\_\{\\tau^\{\\mathbf\{v\}\}\}R\(M\(\\tau^\{\\mathbf\{x\}\},\\tau^\{\\mathbf\{v\}\}\)\)

Algorithm 5๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsfor ligand generation and optimizationRequired Input:๐’Ÿ\\mathop\{\\mathcal\{D\}\}\\limits,๐’ซ\\mathop\{\\mathcal\{P\}\}\\limits, iteration numberNN

1:

n=sampleNumAtomsโ€‹\(๐’Ÿ,๐’ซ\)n=\\text\{\{sampleNumAtoms\}\}\(\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\},\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\)โŠณ\\trianglerightsample the number of ligand atoms from pocket size

2:

\{๐ณi,T๐šก\}i=1nโˆผ๐’ฉโ€‹\(0,๐ˆ\),\{๐ฑTg\}i=1n=\{๐ณT๐šก\}i=1n\\\{\\mathbf\{z\}^\{\\mathtt\{x\}\}\_\{i,T\}\\\}\_\{i=1\}^\{n\}\\sim\\mathcal\{N\}\(0,\\mathbf\{I\}\),\\quad\\\{\\mathbf\{x\}^\{g\}\_\{T\}\\\}\_\{i=1\}^\{n\}=\\\{\\mathbf\{z\}^\{\\mathtt\{x\}\}\_\{T\}\\\}\_\{i=1\}^\{n\}โŠณ\\trianglerightinitialize positions ofnnligand atoms

3:

\{๐ณi,T๐šŸ\}i=1nโˆผGumbelโ€‹\(0,1\)K,\{๐ฏTg\}i=1n=\{onehotโ€‹\(argโกmaxkโˆˆ\[K\]โก๐ณi,T,k๐šŸ\)\}i=1n\\\{\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{i,T\}\\\}\_\{i=1\}^\{n\}\\sim\\mathrm\{Gumbel\}\(0,1\)^\{K\},\\quad\\\{\\mathbf\{v\}^\{g\}\_\{T\}\\\}\_\{i=1\}^\{n\}=\\left\\\{\\mathrm\{onehot\}\\\!\\left\(\\arg\\max\_\{k\\in\[K\]\}\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{i,T,k\}\\right\)\\right\\\}\_\{i=1\}^\{n\}โŠณ\\trianglerightinitialize types ofnnligand atoms

4:

ฯ„๐ฑโ†โˆ…,ฯ„๐ฏโ†โˆ…\\tau^\{\\mathbf\{x\}\}\\leftarrow\\emptyset,\\;\\tau^\{\\mathbf\{v\}\}\\leftarrow\\emptysetโŠณ\\trianglerightinitialize noise trajectories

5:

๐ฌp,โ„‹p=๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\(๐ฑp,๐ฏp,๐ฑr,๐ฏr\)\\mathbf\{s\}^\{p\},\\mathcal\{H\}^\{p\}=\\mbox\{$\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits$\}\(\\mathbf\{x\}^\{p\},\\mathbf\{v\}^\{p\},\\mathbf\{x\}^\{r\},\\mathbf\{v\}^\{r\}\)โŠณ\\trianglerightencode pocket into embeddings using๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits

6:for

t=Tt=Tto

11do

7:

\(๐ฑ~0,tg,๐ฑ~0,tg\)=๐—‰๐–ผ๐–ซ๐–ฆ\(๐ฑtg,๐ฏtg,๐ฌp,โ„‹p\)\(\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{0,t\},\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{0,t\}\)=\\mbox\{$\\mathop\{\\mathsf\{pcLG\}\}\\limits$\}\(\\mathbf\{x\}^\{g\}\_\{t\},\\mathbf\{v\}^\{g\}\_\{t\},\\mathbf\{s\}^\{p\},\\mathcal\{H\}^\{p\}\)โŠณ\\trianglerightpredict noise\-free ligand using the๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limits

8:

๐ณtโˆ’1๐šกโˆผ๐’ฉโ€‹\(0,๐ˆ\),๐ฑtโˆ’1g=qโ€‹\(๐ฑtโˆ’1g\|๐ฑtg,๐ฑ~0,tg\)\\mathbf\{z\}^\{\\mathtt\{x\}\}\_\{t\-1\}\\sim\\mathcal\{N\}\(0,\\mathbf\{I\}\),\\quad\\mathbf\{x\}^\{g\}\_\{t\-1\}=q\(\\mathbf\{x\}^\{g\}\_\{t\-1\}\|\\mathbf\{x\}^\{g\}\_\{t\},\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{0,t\}\)โŠณ\\trianglerightsample๐ฑtโˆ’1g\\mathbf\{x\}^\{g\}\_\{t\-1\}using Gaussian posterior \(Equation[16](https://arxiv.org/html/2607.12349#Sx6.E16)\)

9:

๐ณtโˆ’1๐šŸโˆผGumbelโ€‹\(0,1\),๐ฏtโˆ’1g=qโ€‹\(๐ฏtโˆ’1g\|๐ฏtg,๐ฑ~0,tg\)\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{t\-1\}\\sim\\text\{Gumbel\}\(0,1\),\\quad\\mathbf\{v\}^\{g\}\_\{t\-1\}=q\(\\mathbf\{v\}^\{g\}\_\{t\-1\}\|\\mathbf\{v\}^\{g\}\_\{t\},\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{0,t\}\)โŠณ\\trianglerightsample๐ฏtโˆ’1g\\mathbf\{v\}^\{g\}\_\{t\-1\}using categorical posterior \(Equation[17](https://arxiv.org/html/2607.12349#Sx6.E17)\)

10:

ฯ„๐ฑโ†ฯ„๐ฑโˆช\{๐ณtโˆ’1๐šก\},ฯ„๐ฏโ†ฯ„๐ฏโˆช\{๐ณtโˆ’1๐šŸ\}\\tau^\{\\mathbf\{x\}\}\\leftarrow\\tau^\{\\mathbf\{x\}\}\\cup\\\{\\mathbf\{z\}^\{\\mathtt\{x\}\}\_\{t\-1\}\\\},\\;\\;\\tau^\{\\mathbf\{v\}\}\\leftarrow\\tau^\{\\mathbf\{v\}\}\\cup\\\{\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{t\-1\}\\\}โŠณ\\trianglerightappend noise to trajectories

11:endfor

12:

๐’Ÿbest=๐–ญ๐—ˆ๐—‚๐—Œ๐–พ๐–ฎ๐—‰๐—โ€‹\(ฯ„๐ฑ,ฯ„๐ฑ,N\)\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\_\{\\text\{best\}\}=\\mathsf\{NoiseOpt\}\(\\tau^\{\\mathbf\{x\}\},\\tau^\{\\mathbf\{x\}\},N\)โŠณ\\trianglerightAlgorithm[3](https://arxiv.org/html/2607.12349#alg3), noise optimization

13:return

๐’Ÿbest\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\_\{\\text\{best\}\}

The overall๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsframework is summarized in Algorithm[5](https://arxiv.org/html/2607.12349#alg5)\. At a high level,๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsrecords the noise trajectories produced by๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsduring sampling and then optimizes these trajectories to refine the resulting ligand\. Specifically, it first samples the ligand atom count based on the pocket size and initializes atom positions with Gaussian noise and atom types with Gumbel noise\. Conditioned on pocket embeddings produced by๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\-โ€‹๐–พ๐—‡๐–ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits,๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsruns the generation process fromt=Tt=Tto11: at each step,๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitsestimates the noise\-free ligand\(๐ฑ~0,tg,๐ฑ~0,tg\)\(\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{0,t\},\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{0,t\}\), after which atom positions and types are updated by sampling from the corresponding Gaussian and categorical posteriors\. The injected position and type noises throughout sampling are recorded to the trajectories\(ฯ„๐ฑ,ฯ„๐ฏ\)\(\\tau^\{\\mathbf\{x\}\},\\tau^\{\\mathbf\{v\}\}\)\. Finally, the noise optimization module refines the sampled ligand by optimizing the trajectories forNNiterations and returns the optimized molecule๐’Ÿ\\mathop\{\\mathcal\{D\}\}\\limits\.

## Data Availability

## Code Availability

## Acknowledgements

This project was made possible, in part, by support from the National Science Foundation grant nos\. 2435819 \(X\.N\.\) and 2450988 \(M\.H\.\), the National Library of Medicine grant no\. 1R01LM014385 \(X\.N\.\), the National Center for Advancing Translational Sciences grant no\. UM1TR004548 \(X\.N\.\), and Sanofi iDEA\-TECH Awards North America \(X\.N\.\)\. Any opinions, findings, and conclusions or recommendations expressed in this manuscript are those of the authors, and do not necessarily reflect the views of the funding agencies\. We thank Benjamin Burns, Reza Averly, Maggie Samaan, and Trieu Nguyen for their help with paper writing\. We are grateful for their careful review, constructive feedback, and assistance in improving the clarity and organization of the manuscript\. We thank Avery Meyer for her contributions to the design and implementation of the graphical user interface, which helps improve the usability and accessibility of the method\.

## Author Contributions

X\.N\. conceived the research and conducted the project administration\. M\.H\. and X\.N\. investigated the research, obtained funding and resources for the research, and supervised the student authors \(X\.N\. supervised R\.G, Z\.C\., and F\.B\.; M\.H\. supervised J\.P\.\)\. R\.G\., Z\.C\. and X\.N\. designed the computational methodologies of๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limits\. J\.P\. and M\.H\. designed the computational methodologies of๐—‰๐–บ๐–ฎ๐–ฏ๐–ณ\\mathop\{\\mathsf\{paOPT\}\}\\limits\. R\.G\. conducted data curation, formal analysis, computational methodology \(๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limits\) implementation, result analysis, and visualization\. J\.P\. conducted formal analysis, computational methodology \(๐—‰๐–บ๐–ฎ๐–ฏ๐–ณ\\mathop\{\\mathsf\{paOPT\}\}\\limits\) implementation, result analysis, and visualization\. Z\.C\. contributed to the methodology design and implementation \(๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limits\)\. F\.B\. contributed to formal analysis, result analysis, and visualization\. D\.K\. designed the computational chemistry evaluation and experimental validation of the research and contributed to the computational chemistry results analysis and visualization\. J\.K\. and A\.S\. conducted the molecular synthesis experiments and the results analysis\. H\-P\.B\., Y\.L\. and M\.L\. contributed to the computational chemistry evaluation and experimental validation analysis\. L\.I\. conducted biological assays for PD\-L1\. R\.G\., J\.P\., M\.H\., and X\.N\. drafted the original manuscript\. R\.G\., J\.P\., F\.B\., D\.K\., M\.H\., and X\.N\. conducted the manuscript editing and revision\. All authors reviewed the final paper\.

## References

## Appendix AEquivariance and Invariance

### Equivariance

By definition, a functionfโ€‹\(๐—‘\)f\(\\mathsf\{x\}\)is equivariant if all translation and rotation transformations from the special Euclidean group SE\(3\)\[Atz2021\]applied to the input๐—‘โˆˆโ„3\\mathsf\{x\}\\in\\mathbb\{R\}^\{3\}are mirrored accordingly in the output as follows:

fโ€‹\(๐‘โ€‹๐—‘\+๐ญ\)=๐‘โ€‹fโ€‹\(๐—‘\)\+๐ญ,f\(\\mathbf\{R\}\\mathsf\{x\}\+\\mathbf\{t\}\)=\\mathbf\{R\}f\(\\mathsf\{x\}\)\+\\mathbf\{t\},\(35\)where,๐ญโˆˆโ„3\\mathbf\{t\}\\in\\mathbb\{R\}^\{3\}is a translation transformation and๐‘โˆˆโ„3ร—3\\mathbf\{R\}\\in\\mathbb\{R\}^\{3\\times 3\}\(๐‘๐–ณโ€‹๐‘=๐ˆ\\mathbf\{R\}^\{\\mathsf\{T\}\}\\mathbf\{R\}=\\mathbf\{I\}\) is a rotation transformation\.๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitsleverages GVP and VN\-MLP to ensure๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\\mathop\{\\mathsf\{msPRL\}\}\\limitsand๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{pcLG\}\}\\limits$\}\}\\limitsare equivariant, effectively capturing the geometric features of objects regardless of any translation or rotation transformations\.

### Invariance

A functionfโ€‹\(๐—‘\)f\(\\mathsf\{x\}\)is invariant if its output remains constant under all translation and rotation transformations of the input๐—‘\\mathsf\{x\}:

fโ€‹\(๐‘โ€‹๐—‘\+๐ญ\)=fโ€‹\(๐—‘\),f\(\\mathbf\{R\}\\mathsf\{x\}\+\\mathbf\{t\}\)=f\(\\mathsf\{x\}\),\(36\)where๐ญ\\mathbf\{t\}and๐‘\\mathbf\{R\}represents any translation and rotation transformation, respectively\.๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limitslearns invariant scalar embeddings for pocket atoms and ligand atoms, capturing inherent features \(e\.g\., atom features\) that remain invariant to any translation or rotation transformation\.

## Appendix BForward Diffusion

In the forward process,๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{pcLG\}\}\\limits$\}\}\\limitsadds noises step by step to the atom position \(๐ฑi,tg\\mathbf\{x\}^\{g\}\_\{i,t\}\) and atom feature \(๐ฏi,tg\\mathbf\{v\}^\{g\}\_\{i,t\}\) in the training ligands\. For brevity, in this section, we eliminate the subscriptiiin the notations when no ambiguity arises\. The probability of atom positions๐ฑtg\\mathbf\{x\}^\{g\}\_\{t\}sampled given๐ฑtโˆ’1g\\mathbf\{x\}^\{g\}\_\{t\-1\}, denoted asqโ€‹\(๐ฑtg\|๐ฑtโˆ’1g\)q\(\\mathbf\{x\}^\{g\}\_\{t\}\|\\mathbf\{x\}^\{g\}\_\{t\-1\}\), is defined as follows:

qโ€‹\(๐ฑtg\|๐ฑtโˆ’1g\)=๐’ฉโ€‹\(๐ฑtg\|1โˆ’ฮฒt๐ฑโ€‹๐ฑtโˆ’1g,ฮฒt๐ฑโ€‹๐ˆ\),q\(\\mathbf\{x\}^\{g\}\_\{t\}\|\\mathbf\{x\}^\{g\}\_\{t\-1\}\)=\\mathcal\{N\}\(\\mathbf\{x\}^\{g\}\_\{t\}\|\\sqrt\{1\-\\beta^\{\\mathbf\{x\}\}\_\{t\}\}\\mathbf\{x\}^\{g\}\_\{t\-1\},\\beta^\{\\mathbf\{x\}\}\_\{t\}\\mathbf\{I\}\),\(37\)where๐’ฉโ€‹\(โ‹…\)\\mathcal\{N\}\(\\cdot\)is a Gaussian distribution of๐ฑtg\\mathbf\{x\}^\{g\}\_\{t\}with mean1โˆ’ฮฒt๐ฑโ€‹๐ฑtโˆ’1g\\sqrt\{1\-\\beta\_\{t\}^\{\\mathbf\{x\}\}\}\\mathbf\{x\}^\{g\}\_\{t\-1\}and covarianceฮฒt๐ฑโ€‹๐ˆ\\beta\_\{t\}^\{\\mathbf\{x\}\}\\mathbf\{I\}\. The probability of atom features at time steptt,๐ฏtg\\mathbf\{v\}^\{g\}\_\{t\}, given that at time steptโˆ’1t\-1,๐ฏtโˆ’1g\\mathbf\{v\}^\{g\}\_\{t\-1\}is defined as follows:

qโ€‹\(๐ฏtg\|๐ฏtโˆ’1g\)=๐’žโ€‹\(๐ฏtg\|\(1โˆ’ฮฒt๐ฏ\)โ€‹๐ฏtโˆ’1g\+ฮฒt๐ฏโ€‹๐Ÿ/da\),q\(\\mathbf\{v\}^\{g\}\_\{t\}\|\\mathbf\{v\}^\{g\}\_\{t\-1\}\)=\\mathcal\{C\}\(\\mathbf\{v\}^\{g\}\_\{t\}\|\(1\-\\beta^\{\\mathbf\{v\}\}\_\{t\}\)\\mathbf\{v\}^\{g\}\_\{t\-1\}\+\\beta^\{\\mathbf\{v\}\}\_\{t\}\\mathbf\{1\}/d\_\{a\}\),\(38\)where๐’ž\\mathcal\{C\}is a categorical distribution\.

Given the above definitions, the probability of atom positions \(๐ฑtg\\mathbf\{x\}^\{g\}\_\{t\}\) and atom features \(๐ฏtg\\mathbf\{v\}^\{g\}\_\{t\}\) at any time stepttcan be derived from those at the initial time step \(๐ฑ0g\\mathbf\{x\}^\{g\}\_\{0\}and๐ฏ0g\\mathbf\{v\}^\{g\}\_\{0\}\) as follows:

qโ€‹\(๐ฑtg\|๐ฑ0g\)\\displaystyle q\(\\mathbf\{x\}^\{g\}\_\{t\}\|\\mathbf\{x\}^\{g\}\_\{0\}\)=๐’ฉโ€‹\(๐ฑtg\|ฮฑยฏt๐ฑโ€‹๐ฑ0g,\(1โˆ’ฮฑยฏt๐ฑ\)โ€‹๐ˆ\),\\displaystyle=\\mathcal\{N\}\(\\mathbf\{x\}^\{g\}\_\{t\}\|\\sqrt\{\\mbox\{$\\mathop\{\\bar\{\\alpha\}\}\\limits$\}^\{\\mathbf\{x\}\}\_\{t\}\}\\mathbf\{x\}^\{g\}\_\{0\},\(1\-\\mbox\{$\\mathop\{\\bar\{\\alpha\}\}\\limits$\}^\{\\mathbf\{x\}\}\_\{t\}\)\\mathbf\{I\}\),\(39\)qโ€‹\(๐ฏtg\|๐ฏ0g\)\\displaystyle q\(\\mathbf\{v\}^\{g\}\_\{t\}\|\\mathbf\{v\}^\{g\}\_\{0\}\)=๐’žโ€‹\(๐ฏtg\|ฮฑยฏt๐ฏ๐ฏ0g\+\(1โˆ’ฮฑยฏt๐ฏ\)โ€‹๐Ÿ/K\),\\displaystyle=\\mathcal\{C\}\(\\mathbf\{v\}^\{g\}\_\{t\}\|\\mbox\{$\\mathop\{\\bar\{\\alpha\}\}\\limits$\}^\{\\mathbf\{v\}\}\_\{t\}\\mathbf\{v\}^\{g\}\_\{0\}\+\(1\-\\mbox\{$\\mathop\{\\bar\{\\alpha\}\}\\limits$\}^\{\\mathbf\{v\}\}\_\{t\}\)\\mathbf\{1\}/K\),\(40\)whereฮฑยฏt๐šž\\displaystyle\\text\{where \}\\mbox\{$\\mathop\{\\bar\{\\alpha\}\}\\limits$\}^\{\\mathtt\{u\}\}\_\{t\}=โˆฯ„=1tฮฑฯ„๐šž,ฮฑฯ„๐šž=1โˆ’ฮฒฯ„๐šž,๐šž=๐ฑโ€‹orโ€‹๐ฏ,\\displaystyle=\\displaystyle\\prod\_\{\\tau=1\}^\{t\}\\alpha^\{\\mathtt\{u\}\}\_\{\\tau\},\\ \\alpha^\{\\mathtt\{u\}\}\_\{\\tau\}=1\-\\beta^\{\\mathtt\{u\}\}\_\{\\tau\},\\ \{\\mathtt\{u\}\}=\{\\mathbf\{x\}\}\\text\{ or \}\{\\mathbf\{v\}\},\\;\\;\\;\(41\)whereฮฑยฏt๐šž\\bar\{\\alpha\}^\{\\mathtt\{u\}\}\_\{t\}is a weight decreasing monotonically from 1 to 0 overt=\[1,T\]t=\[1,T\]\. Specifically,ฮฑยฏt๐šž\\mbox\{$\\mathop\{\\bar\{\\alpha\}\}\\limits$\}^\{\\mathtt\{u\}\}\_\{t\}\(๐šž=๐ฑโ€‹orโ€‹๐ฏ\\mathtt\{u\}=\{\\mathbf\{x\}\}\\text\{ or \}\{\\mathbf\{v\}\}\) approaches 1 astโ†’1t\\rightarrow 1, allowing๐ฑtg\\mathbf\{x\}^\{g\}\_\{t\}or๐ฏtg\\mathbf\{v\}^\{g\}\_\{t\}to approximate๐ฑ0g\\mathbf\{x\}^\{g\}\_\{0\}or๐ฏ0g\\mathbf\{v\}^\{g\}\_\{0\}\. Conversely,ฮฑยฏt๐šž\\mbox\{$\\mathop\{\\bar\{\\alpha\}\}\\limits$\}^\{\\mathtt\{u\}\}\_\{t\}\(๐šž=๐ฑโ€‹orโ€‹๐ฏ\\mathtt\{u\}=\{\\mathbf\{x\}\}\\text\{ or \}\{\\mathbf\{v\}\}\) approaches 0 astโ†’Tt\\rightarrow T, which makesqโ€‹\(xTgโˆฃx0g\)q\(x\_\{T\}^\{g\}\\mid x\_\{0\}^\{g\}\)resemble๐’ฉโ€‹\(0,I\)\\mathcal\{N\}\(0,I\)andqโ€‹\(vTgโˆฃv0g\)q\(v\_\{T\}^\{g\}\\mid v\_\{0\}^\{g\}\)resemble๐’žโ€‹\(1/da\)\\mathcal\{C\}\(1/d\_\{a\}\)\.

As shown by Ho*et al\.*\[ho2020ddpm\], the ground\-truth Normal posterior of atom positions,pโ€‹\(๐ฑtโˆ’1g\|๐ฑtg,๐ฑ0g\)p\(\\mathbf\{x\}^\{g\}\_\{t\-1\}\|\\mathbf\{x\}^\{g\}\_\{t\},\\mathbf\{x\}^\{g\}\_\{0\}\), could be calculated in a closed form as below:

qโ€‹\(๐ฑtโˆ’1g\|๐ฑtg,๐ฑ0g\)=๐’ฉโ€‹\(๐ฑtโˆ’1g\|ฮผโ€‹\(๐ฑtg,๐ฑ0g\),ฮฒ~t๐ฑโ€‹๐ˆ\),\\displaystyle q\(\\mathbf\{x\}^\{g\}\_\{t\-1\}\|\\mathbf\{x\}^\{g\}\_\{t\},\\mathbf\{x\}^\{g\}\_\{0\}\)=\\mathcal\{N\}\(\\mathbf\{x\}^\{g\}\_\{t\-1\}\|\\mu\(\\mathbf\{x\}^\{g\}\_\{t\},\\mathbf\{x\}^\{g\}\_\{0\}\),\\tilde\{\\beta\}^\{\\mathbf\{x\}\}\_\{t\}\\mathbf\{I\}\),\(42\)ฮผโ€‹\(๐ฑtg,๐ฑ0g\)=ฮฑยฏtโˆ’1๐ฑโ€‹ฮฒt๐ฑ1โˆ’ฮฑยฏt๐ฑโ€‹๐ฑ0g\+ฮฑt๐ฑโ€‹\(1โˆ’ฮฑยฏtโˆ’1๐ฑ\)1โˆ’ฮฑยฏt๐ฑโ€‹๐ฑtg,\\displaystyle\\mu\(\\mathbf\{x\}^\{g\}\_\{t\},\\mathbf\{x\}^\{g\}\_\{0\}\)=\\frac\{\\sqrt\{\\bar\{\\alpha\}^\{\\mathbf\{x\}\}\_\{t\-1\}\}\\beta^\{\\mathbf\{x\}\}\_\{t\}\}\{1\-\\bar\{\\alpha\}^\{\\mathbf\{x\}\}\_\{t\}\}\\mathbf\{x\}^\{g\}\_\{0\}\\\!\+\\\!\\frac\{\\sqrt\{\\alpha^\{\\mathbf\{x\}\}\_\{t\}\}\(1\-\\bar\{\\alpha\}^\{\\mathbf\{x\}\}\_\{t\-1\}\)\}\{1\-\\bar\{\\alpha\}^\{\\mathbf\{x\}\}\_\{t\}\}\\mathbf\{x\}^\{g\}\_\{t\},\(43\)ฮฒ~t๐ฑ=1โˆ’ฮฑยฏtโˆ’1๐ฑ1โˆ’ฮฑยฏt๐ฑโ€‹ฮฒt๐ฑ\.\\displaystyle\\tilde\{\\beta\}^\{\\mathbf\{x\}\}\_\{t\}=\\frac\{1\-\\bar\{\\alpha\}^\{\\mathbf\{x\}\}\_\{t\-1\}\}\{1\-\\bar\{\\alpha\}^\{\\mathbf\{x\}\}\_\{t\}\}\\beta^\{\\mathbf\{x\}\}\_\{t\}\.\\;\\;\\;\(44\)Similarly, as shown in Hoogeboom*et al\.*\[hoogeboom22diff\], the ground\-truth categorical posterior of atom featurespโ€‹\(๐ฏtโˆ’1g\|๐ฏtg,๐ฏ0g\)p\(\\mathbf\{v\}^\{g\}\_\{t\-1\}\|\\mathbf\{v\}^\{g\}\_\{t\},\\mathbf\{v\}^\{g\}\_\{0\}\)can be calculated as below:

qโ€‹\(๐ฏtโˆ’1g\|๐ฏtg,๐ฏ0g\)=๐’žโ€‹\(๐ฏtโˆ’1g\|๐œโ€‹\(๐ฏtg,๐ฏ0g\)\),\\displaystyle q\(\\mathbf\{v\}^\{g\}\_\{t\-1\}\|\\mathbf\{v\}^\{g\}\_\{t\},\\mathbf\{v\}^\{g\}\_\{0\}\)=\\mathcal\{C\}\(\\mathbf\{v\}^\{g\}\_\{t\-1\}\|\\mathbf\{c\}\(\\mathbf\{v\}^\{g\}\_\{t\},\\mathbf\{v\}^\{g\}\_\{0\}\)\),\(45\)๐œโ€‹\(๐ฏtg,๐ฏ0g\)=๐œ~/โˆ‘k=1Kc~k,\\displaystyle\\mathbf\{c\}\(\\mathbf\{v\}^\{g\}\_\{t\},\\mathbf\{v\}^\{g\}\_\{0\}\)=\\tilde\{\\mathbf\{c\}\}/\{\\sum\_\{k=1\}^\{K\}\\tilde\{c\}\_\{k\}\},\(46\)๐œ~=\[ฮฑt๐ฏโ€‹๐ฏtg\+1โˆ’ฮฑt๐ฏda\]โŠ™\[ฮฑยฏtโˆ’1๐ฏโ€‹๐ฏ0g\+1โˆ’ฮฑยฏtโˆ’1๐ฏda\],\\displaystyle\\tilde\{\\mathbf\{c\}\}=\[\\alpha^\{\\mathbf\{v\}\}\_\{t\}\\mathbf\{v\}^\{g\}\_\{t\}\+\\frac\{1\-\\alpha^\{\\mathbf\{v\}\}\_\{t\}\}\{d\_\{a\}\}\]\\odot\[\\bar\{\\alpha\}^\{\\mathbf\{v\}\}\_\{t\-1\}\\mathbf\{v\}^\{g\}\_\{0\}\+\\frac\{1\-\\bar\{\\alpha\}^\{\\mathbf\{v\}\}\_\{t\-1\}\}\{d\_\{a\}\}\],\(47\)where๐œโ€‹\(๐ฏtg,๐ฏ0g\)\\mathbf\{c\}\(\\mathbf\{v\}^\{g\}\_\{t\},\\mathbf\{v\}^\{g\}\_\{0\}\)denotes the probability over thedad\_\{a\}classes,c~k\\tilde\{c\}\_\{k\}denotes the likelihood of thekk\-th class, andโŠ™\\odotis the element\-wise product operation\.

## Appendix CBackward Generative Process

In the backward process,๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{pcLG\}\}\\limits$\}\}\\limitsgenerates realistic binding ligands from random noise\. Particularly, conditioned on๐ฑi,tg\\mathbf\{x\}^\{g\}\_\{i,t\}and๐—‘~i,0,tg\\tilde\{\\mathsf\{x\}\}^\{g\}\_\{i,0,t\}, the probabilitypโ€‹\(๐ฑi,tโˆ’1g\|๐ฑi,tg\)p\(\\mathbf\{x\}^\{g\}\_\{i,t\-1\}\|\\mathbf\{x\}^\{g\}\_\{i,t\}\)could be estimated using the approximated posteriorp๐šฏโ€‹\(๐ฑi,tโˆ’1g\|๐ฑi,tg,๐—‘~i,0,tg\)p\_\{\\boldsymbol\{\\Theta\}\}\(\\mathbf\{x\}^\{g\}\_\{i,t\-1\}\|\\mathbf\{x\}^\{g\}\_\{i,t\},\\tilde\{\\mathsf\{x\}\}^\{g\}\_\{i,0,t\}\), as shown in Ho*et al\.*\[ho2020ddpm\]\. Same as Appendix[B](https://arxiv.org/html/2607.12349#A2), we eliminate the subscriptiiin the notations when no ambiguity arises\. The approximation ofpโ€‹\(๐ฑtโˆ’1g\|๐ฑtg\)p\(\\mathbf\{x\}^\{g\}\_\{t\-1\}\|\\mathbf\{x\}^\{g\}\_\{t\}\)is calculated as follows:

pโ€‹\(๐ฑtโˆ’1g\|๐ฑtg\)\\displaystyle p\(\\mathbf\{x\}^\{g\}\_\{t\-1\}\|\\mathbf\{x\}^\{g\}\_\{t\}\)โ‰ˆqโ€‹\(๐ฑtโˆ’1g\|๐ฑtg,๐ฑ~0,tg\)\\displaystyle\\approx q\(\\mathbf\{x\}^\{g\}\_\{t\-1\}\|\\mathbf\{x\}^\{g\}\_\{t\},\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{0,t\}\)\(48\)=๐’ฉโ€‹\(๐ฑtโˆ’1g\|ฮผโ€‹\(๐ฑtg,๐ฑ~0,tg\),ฮฒ~t๐ฑโ€‹๐ˆ\),\\displaystyle=\\mathcal\{N\}\(\\mathbf\{x\}^\{g\}\_\{t\-1\}\|\\mu\(\\mathbf\{x\}^\{g\}\_\{t\},\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{0,t\}\),\\tilde\{\\beta\}\_\{t\}^\{\\mathbf\{x\}\}\\mathbf\{I\}\),whereฮผโ€‹\(๐ฑtg,๐—‘~0,tg\)\\mu\(\\mathbf\{x\}^\{g\}\_\{t\},\\tilde\{\\mathsf\{x\}\}^\{g\}\_\{0,t\}\)is an estimate ofฮผโ€‹\(๐ฑtg,๐ฑ0g\)\\mu\(\\mathbf\{x\}^\{g\}\_\{t\},\\mathbf\{x\}^\{g\}\_\{0\}\)by replacing๐ฑ0g\\mathbf\{x\}^\{g\}\_\{0\}with its approximation๐—‘~0,tg\\tilde\{\\mathsf\{x\}\}^\{g\}\_\{0,t\}in Equation[42](https://arxiv.org/html/2607.12349#A2.E42)\. Similarly, as shown in Hoogeboom\[hoogeboom22diff\], given๐ฏtg\\mathbf\{v\}^\{g\}\_\{t\}and๐—~0,tg\\tilde\{\\mathsf\{v\}\}^\{g\}\_\{0,t\}, the probability of๐ฏtโˆ’1g\\mathbf\{v\}^\{g\}\_\{t\-1\}conditioned on๐ฏtg\\mathbf\{v\}^\{g\}\_\{t\},pโ€‹\(๐ฏtโˆ’1g\|๐ฏtg\)p\(\\mathbf\{v\}^\{g\}\_\{t\-1\}\|\\mathbf\{v\}^\{g\}\_\{t\}\), can be estimated by the approximated posteriorqโ€‹\(๐ฏtโˆ’1g\|๐ฏtg,๐—~0,tg\)q\(\\mathbf\{v\}^\{g\}\_\{t\-1\}\|\\mathbf\{v\}^\{g\}\_\{t\},\\tilde\{\\mathsf\{v\}\}^\{g\}\_\{0,t\}\)as below:

pโ€‹\(๐ฏtโˆ’1g\|๐ฏtg\)โ‰ˆqโ€‹\(๐ฏtโˆ’1g\|๐ฏtg,๐—~0,tg\)=๐’žโ€‹\(๐ฏtโˆ’1g\|๐œโ€‹\(๐ฏtg,๐—~0,tg\)\),\\displaystyle p\(\\mathbf\{v\}^\{g\}\_\{t\-1\}\|\\mathbf\{v\}^\{g\}\_\{t\}\)\\approx q\(\\mathbf\{v\}^\{g\}\_\{t\-1\}\|\\mathbf\{v\}^\{g\}\_\{t\},\\tilde\{\\mathsf\{v\}\}^\{g\}\_\{0,t\}\)=\\mathcal\{C\}\(\\mathbf\{v\}^\{g\}\_\{t\-1\}\|\\mathbf\{c\}\(\\mathbf\{v\}^\{g\}\_\{t\},\\tilde\{\\mathsf\{v\}\}^\{g\}\_\{0,t\}\)\),\(49\)where๐œโ€‹\(๐ฏtg,๐—~0,tg\)\\mathbf\{c\}\(\\mathbf\{v\}^\{g\}\_\{t\},\\tilde\{\\mathsf\{v\}\}^\{g\}\_\{0,t\}\)is an estimate of๐œโ€‹\(๐ฏtg,๐ฏ0g\)\\mathbf\{c\}\(\\mathbf\{v\}^\{g\}\_\{t\},\\mathbf\{v\}^\{g\}\_\{0\}\)by replacing๐ฏ0g\\mathbf\{v\}^\{g\}\_\{0\}with its estimate๐—~0,tg\\tilde\{\\mathsf\{v\}\}^\{g\}\_\{0,t\}in Equation[45](https://arxiv.org/html/2607.12349#A2.E45)\.

## Appendix DParameters for Reproducibility

In๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\\mathop\{\\mathsf\{conDitar\}\}\\limits, we trained two models๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\\mathop\{\\mathsf\{msPRL\}\}\\limitsand๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{pcLG\}\}\\limits$\}\}\\limitsfor pocket representation learning and pocket\-conditioned ligand generation, respectively\. We implemented both models using Python 3\.9\.18 and PyTorch 2\.1\.0, together with the corresponding PyTorch Geometric dependencies, including torch\-scatter 2\.1\.2, torch\-cluster 1\.6\.3, and torch\-geometric 2\.6\.1\. The environment is built with CUDA 12\.1 \(pytorch\-cuda 12\.1\)\. We trained both models on an NVIDIA A100 GPU with 40GB memory and a CPU with 80GB memory\.

### Parameters formsPRL\\mathop\{\\mathsf\{msPRL\}\}\\limits

In๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\\mathop\{\\mathsf\{msPRL\}\}\\limits, we set the dimension of all the hidden layers, including GVP layers \(Equation[3](https://arxiv.org/html/2607.12349#Sx6.E3)and[4](https://arxiv.org/html/2607.12349#Sx6.E4)\) and MLP layers \(Equation[5](https://arxiv.org/html/2607.12349#Sx6.E5)to[9](https://arxiv.org/html/2607.12349#Sx6.E9)\) as 128, and the dimension of both residue scalar and vector embeddings \(๐ฌr\\mathbf\{s\}^\{r\}andโ„‹r\\mathcal\{H\}^\{r\}\) and pocket atom scalar and vector embeddings \(๐ฌp\\mathbf\{s\}^\{p\}andโ„‹p\\mathcal\{H\}^\{p\}\) as 128\. To represent the pocket structure, we constructed two graphs using thekk\-nearest neighbors based on Euclidean distance: one for pocket atoms withka=16k\_\{a\}=16and one for residues withkr=8k\_\{r\}=8\. We set the layer number of graph neural networks for both atom and residue as 3\. We optimized the๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\\mathop\{\\mathsf\{msPRL\}\}\\limitsmodel with Adam with its parameters \(0\.950, 0\.999\), learning rate 0\.001, and batch size 64\. We trained๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\\mathop\{\\mathsf\{msPRL\}\}\\limitsfor maximum 100 epochs and the training tookโˆผ\\sim36 hours in total\.

### Parameters forpcLG\\mathop\{\\mathsf\{pcLG\}\}\\limits

In๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limits, we set the dimension of all the scalar hidden layers, including GVP layers \(Equation[23](https://arxiv.org/html/2607.12349#Sx6.E23),[25](https://arxiv.org/html/2607.12349#Sx6.E25)and[31](https://arxiv.org/html/2607.12349#Sx6.E31)\) and VN\-MLP and MLP layers \(Equation[27](https://arxiv.org/html/2607.12349#Sx6.E27)to[33](https://arxiv.org/html/2607.12349#Sx6.E33)\) as 128\. We set the dimensions of all the vector hidden layers in GVPs as 32\. We set the number of layersLLin๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitsas 10\. We built the distance\-based atomic graphs for ligands using themm\-nearest neighbors based on Euclidean distance withm=8m=8\. For each ligand atom, we consider its nearestna=32n\_\{a\}=32pocket atoms for protein\-ligand interaction learning\. In addition, we consider residue\-level interaction by connecting each ligand atom to itsnr=4n\_\{r\}=4nearest pocket residues\.

In the forward process of the diffusion model, following Guan*et al\.*\[guan2023targetdiff\], we used a sigmoidฮฒ\\betaschedule for the variance scheduleฮฒt๐ฑ\\beta\_\{t\}^\{\\mathbf\{x\}\}of atom positions to add noises into atom positions as below:

ฮฒt๐ฑ=sigmoidโ€‹\(w1โ€‹\(2โ€‹t/Tโˆ’1\)\)โ€‹\(w2โˆ’w3\)\+w3\\beta\_\{t\}^\{\\mathbf\{x\}\}=\\text\{sigmoid\}\(w\_\{1\}\(2t/T\-1\)\)\(w\_\{2\}\-w\_\{3\}\)\+w\_\{3\}\(50\)in whichwiw\_\{i\}\(ii=1,2, or 3\) withw1=6w\_\{1\}=6,w2=1\.eโˆ’7w\_\{2\}=1\.e\-7andw3=0\.01w\_\{3\}=0\.01are hyperparameters;T=1,000T=1,000is the maximum step\. For atom types, we used a cosineฮฒ\\betaschedule\[nichol2021\]forฮฒt๐ฏ\\beta\_\{t\}^\{\\mathbf\{v\}\}as below:

ฮฑยฏt๐ฏ=fโ€‹\(t\)fโ€‹\(0\),f\(t\)=cos\(t/T\+s1\+sโ‹…ฯ€2\)2\\displaystyle\\bar\{\\alpha\}\_\{t\}^\{\\mathbf\{v\}\}=\\frac\{f\(t\)\}\{f\(0\)\},f\(t\)=\\cos\(\\frac\{t/T\+s\}\{1\+s\}\\cdot\\frac\{\\pi\}\{2\}\)^\{2\}\(51\)ฮฒt๐ฏ=1โˆ’ฮฑt๐ฏ=1โˆ’ฮฑยฏt๐ฏฮฑยฏtโˆ’1๐ฏ\\displaystyle\\beta\_\{t\}^\{\\mathbf\{v\}\}=1\-\\alpha\_\{t\}^\{\\mathbf\{v\}\}=1\-\\frac\{\\bar\{\\alpha\}\_\{t\}^\{\\mathbf\{v\}\}\}\{\\bar\{\\alpha\}\_\{t\-1\}^\{\\mathbf\{v\}\}\}in whichssis a hyperparameter and set as 0\.01\. Same as๐—†๐—Œ๐–ฏ๐–ฑ๐–ซ\\mathop\{\\mathsf\{msPRL\}\}\\limits, we optimized๐—‰๐–ผ๐–ซ๐–ฆ\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{pcLG\}\}\\limits$\}\}\\limitsusing Adam with its parameters \(0\.950, 0\.999\), the learning rate 0\.001, and batch size 16\. The training takesโˆผ\\sim100 hours in total\.

### Parameters forconDitar\-โ€‹dev\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits

In๐–ผ๐—ˆ๐—‡๐–ฃ๐—‚๐—๐–บ๐—‹\-โ€‹๐–ฝ๐–พ๐—\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits, we set the number of optimization iterations toN=10N=10and the step size toฮฑ=0\.1\\alpha=0\.1for noise optimization\. For zero*th*\-order gradient estimation, we useH=4H=4random perturbations with a smoothing parameter ofฮผ=0\.03\\mu=0\.03\.

## Appendix EDetails of๐–ข๐–ฃ๐–ง\\mathop\{\\mathsf\{CDH\}\}\\limitsdataset

Table A1:Summary of๐–ข๐–ฃ๐–ง\\mathop\{\\mathsf\{CDH\}\}\\limits

Similar Articles

TD3B: Transition-Directed Discrete Diffusion for Allosteric Binder Generation

Hugging Face Daily Papers

TD3B is a sequence-based generative framework for designing allosteric binders with specific agonist or antagonist behaviors using transition-directed discrete diffusion. The paper introduces a method to control directional transitions in protein states, addressing limitations of static structure-based design.

Controllable Molecular Generative Foundation Models

arXiv cs.LG

Proposes CoMole, a controllable molecular generative foundation model using motif-aware graph diffusion and reinforcement learning, achieving superior controllability across materials and drug discovery benchmarks.