Generating Developable 3D Molecules via Pocket-Conditioned Diffusion and Property-Aware Optimization
Summary
This paper introduces a novel diffusion-based generative model for structure-based drug design that decouples pocket and ligand representation learning and incorporates multi-scale interaction signals and property-aware optimization to generate developable 3D molecules with improved binding affinity and ADMET properties.
View Cached Full Text
Cached at: 07/15/26, 04:18 AM
# Generating Developable 3D Molecules via Pocket-Conditioned Diffusion and Property-Aware Optimization
Source: [https://arxiv.org/html/2607.12349](https://arxiv.org/html/2607.12349)
Ruoxi Gao1, Jiangweizhi Peng2, Ziqi Chen3, Frazier N\. Baker1, David C\. Kombo4, John L\. Kane, Jr\.4, Andrew A\. Scholte4, Yi Li4, Matthew J\. LaMarche4, Luigi I\. Iconaru5, Hans\-Peter Biemann5, Mingyi Hong6, Xia Ning1,7,8,9 ๐
## Introduction
Drug discovery and development is complex, time\-consuming, and resource\-intensive โ new drugs typically take 12\-15 years to develop\[singh2023drug\]at costs of $378 million to $1\.76 billion\[sertkaya2024costs\]\. To accelerate this process and reduce costs, computational methods, particularly recent generative models \(e\.g\., diffusion\[hoogeboom22diff,guan2023decompdiff\], variational autoencoders\[jin18jtvae\], large language models\[Yu2024\]\), have been employed for*de novo*drug design\. Rather than conducting expensive searches over large libraries to identify potential binding molecules, these models directly generate molecules*in silico*based on the chemical knowledge learned from a vast amount of experimental data\. These models generally follow one of the two conventional drug design frameworks:\(1\)structure\-based drug design \(SBDD\), where molecules are designed to fit into the binding pocket structure of a given target; and\(2\)ligand\-based drug design \(LBDD\), where molecules are designed to resemble a known binding ligand\. Compared to LBDD, SBDD offers a more direct and mechanistic entry point to molecular discovery by explicitly exploiting pocket\-ligand interactions, thereby providing a principled starting point for subsequent ligand\-based modeling and optimization\. Recent development of generative models for SBDD\[chen2025generating,guan2023decompdiff,guan2023targetdiff\]features diffusion\[ho2020ddpm\], a process that gradually transforms noise into molecular structures, as the leading paradigm\. By learning the underlying distributions of chemically and structurally feasible molecules from data and leveraging the distributions to generate novel molecules, diffusion holds substantial promise to revolutionize*de novo*drug design by rapidly producing high\-quality, functionally relevant molecular candidates\.
Still, existing diffusion\-based models for SBDD leave substantial room for improvement\. Typically, diffusion\-based SBDD models tend to couple pocket and ligand representation learning, using a shared representation to reflect the complex pocket\-ligand interactions\. However, such coupling could amplify early generation errors, leading to progressively degraded representations during diffusionโs iterative refinement process\. Instead, dedicated pocket representation learning may better capture binding pocket structures and properties, and thus, ligand generation can be conditioned on robust pocket features that favor strong binding\. Furthermore, many existing models do not yet fully leverage interaction signals across multiple scales โ ranging from atom\-level to residue\-level โ in modeling pocket\-ligand interactions\. Incorporating rich, multi\-scale interaction signals could allow models to better represent the biochemical environment of the binding site, capturing both local atomic interactions and broader residue\-level context that jointly govern ligand recognition and binding\. In addition, generative SBDD models typically prioritize the generation of ligands with optimized binding affinity, leaving other critical drug properties, such as Absorption, Distribution, Metabolism, Excretion, and Toxicity \(ADMET\), to be addressed in downstream lead optimization phase\. While this aligns with the conventional drug development process\[hughes2011principles\], it also underscores the opportunity for a more proactive strategy that considers a broader set of drug\-relevant properties at the stage of initial molecule generation, thereby increasing the likelihood of ultimate success\. These insights motivate the design of a new, decoupled, multi\-scale, and developability\-aware diffusion\-based framework for SBDD\.
Here, we introduce๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits, a conditional diffusion\-based SBDD framework capable of generating ligands with strong binding affinities and favorable ADMET properties\.๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsconsists of three modules:\(1\)a pretrainedmulti\-scalepocketrepresentationlearning module, referred to as๐๐๐ฏ๐ฑ๐ซ\\mathop\{\\mathsf\{msPRL\}\}\\limits, which encodes the binding pocketโs atomic and residue\-level composition and structure into expressive representations that will guide ligand generation;\(2\)a pocket\-conditioneddiffusion model that generates ligands for strongtarget binding, referred to as๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limits; and\(3\)a generation\-time,property\-awareoptimization mechanism for drug developability built upon๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limits, referred to as๐๐บ๐ฎ๐ฏ๐ณ\\mathop\{\\mathsf\{paOPT\}\}\\limits\.๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsemploys apocket\-conditionedligandgenerator, referred to as๐๐ผ๐ซ๐ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limits, to predict atom types and structures of a potential binding ligand from the pocket representation learned from๐๐๐ฏ๐ฑ๐ซ\\mathop\{\\mathsf\{msPRL\}\}\\limitsand pocket\-ligand interactions at both the atom and residue levels\. The๐๐ผ๐ซ๐ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitsโs predicted atom positions and types are used to direct the diffusion process of๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitstowards the final generated ligands of high binding affinities\. Beyond binding affinity, the training\-free, plug\-and\-play๐๐บ๐ฎ๐ฏ๐ณ\\mathop\{\\mathsf\{paOPT\}\}\\limitssteers ligand generation towards both high binding affinities and favorable developability in terms of ADMET profiles\. In summary,๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitspresents a new diffusion\-based paradigm for generating ligands with high binding affinity and favorable drug developability profiles, bridging binding\-driven design and developability optimization in real\-world drug discovery and development settings\. Fig\.[1](https://arxiv.org/html/2607.12349#Sx1.F1)presents the overview architecture of๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\.
To evaluate๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits, as well as other SBDD methods in the literature\[luo2021sbdd,peng22pocket2mol,guan2023targetdiff,guan2023decompdiff\], we carefully curate a new dataset, denoted as๐ข๐ฃ๐ง\\mathop\{\\mathsf\{CDH\}\}\\limits, on top of the widely\-used benchmark dataset CrossDocked2020\[Francoeur2020\], denoted as๐ข๐ฃ\\mathop\{\\mathsf\{CD\}\}\\limits\. While well adopted,๐ข๐ฃ\\mathop\{\\mathsf\{CD\}\}\\limitsincludes both human and non\-human protein targets \(33% and 67%, respectively\)\. This poses risks to translational efficacy and safety, as even for conserved targets, sequence, structural, dynamics, and conformational differences between human and non\-human proteins can alter ligand binding and target engagement\[marshall2023poor\]\. To address this issue,๐ข๐ฃ๐ง\\mathop\{\\mathsf\{CDH\}\}\\limitsis curated to consist exclusively of human targets, including human targets from๐ข๐ฃ\\mathop\{\\mathsf\{CD\}\}\\limits, the human orthologs of non\-human targets in๐ข๐ฃ\\mathop\{\\mathsf\{CD\}\}\\limits, and some additional, carefully\-selected, experimentally validated human targets from the Protein Database Bank \(PDB\)\[berman2000protein\]that are of high therapeutic interest in life\-threatening conditions \(e\.g\., cancers\)\. Over๐ข๐ฃ๐ง\\mathop\{\\mathsf\{CDH\}\}\\limitsand๐ข๐ฃ\\mathop\{\\mathsf\{CD\}\}\\limits, we extensively compare๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsagainst seven state\-of\-the\-art baselines in generating high\-quality ligands\. Computational results demonstrate that๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsgenerates ligands with superior predicted binding affinity, achieving average scores ofโ8\.85\-8\.85kcal/mol and a 7\.4% improvement over the best baseline on๐ข๐ฃ๐ง\\mathop\{\\mathsf\{CDH\}\}\\limits, and๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsfurther improves ADMET properties on๐ข๐ฃ๐ง\\mathop\{\\mathsf\{CDH\}\}\\limitswith an average improvement of up to 73% across five ADMET properties while maintaining comparable predicted binding affinities\.
To further demonstrate that๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitscan propose developable molecules, we conduct case studies on the programmed death\-ligand 1 \(PD\-L1\), and colony\-stimulating factor\-1 receptor \(CSF1R\) proteins, which are two therapeutically important and validated targets with available structures\. Top\-ranked๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated ligands and their analogs are synthesized and tested in biological assays\. Across the two targets, the selected molecules show promising biological activities\. In particular, two ligands generated by๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsfor PD\-L1 show SPR\-derived binding activities in the low micromolar range, with KD\{\}\_\{\\text\{D\}\}values below 3\.80ฮผ\\muM\. For CSF1R, two analogs derived from๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated ligands show strong kinase activities and high selectivity, with IC50values below 0\.31ฮผ\\muM and promiscuity hit rates of at most 2\.4%\. These biological testing results suggest that๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsis not limited to generating ligands with computationally predicted binding, but can also produce biologically active and developable molecules*in vitro*\.
In summary, we highlight the advantages of๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsas follows:
- โข๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsunifies pocket\-conditioned generation and property optimization within a single framework, incorporating developability considerations directly into the initial drug design process\.
- โข๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsproduces ligands with strong binding affinities leveraging decoupled pocket representation learning\. This decoupling allows๐๐๐ฏ๐ฑ๐ซ\\mathop\{\\mathsf\{msPRL\}\}\\limitsto better capture binding pocket information, thus providing๐๐ผ๐ซ๐ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitswith robust pocket conditioning\.
- โข๐๐๐ฏ๐ฑ๐ซ\\mathop\{\\mathsf\{msPRL\}\}\\limitsand๐๐ผ๐ซ๐ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitsencode multi\-scale features of the binding pocket and pocket\-ligand interactions, allowing molecule generation to be informed by both local atomic details and broader residue\-level context that jointly govern ligand recognition and binding\.
- โข๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsachieves favorable ADMET profiles by incorporating the training\-free๐๐บ๐ฎ๐ฏ๐ณ\\mathop\{\\mathsf\{paOPT\}\}\\limitsto directly optimize those properties during the generation process of๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limits\. This allows external ADMET signals and pocket conditioning to jointly guide ligand generation\.
- โขCase studies with extensive*in silico*analyses on PD\-L1 and CSF1R demonstrate that๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitscan generate developable molecules with binding modes similar to those of known binding ligands of these two targets\.
- โขFifteen top\-ranked๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-designed molecules for PD\-L1 and CSF1R, and their derivatives and analogs, have been experimentally synthesized and tested\. Two hits have been identified with target inhibition at nanomolar concentrations, and seven hits have target inhibition at low micromolar concentrations\.
Fig\. 1:The overall schematic diagram of๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\.a,Pocket representation pretraining,๐๐๐ฏ๐ฑ๐ซ\\mathop\{\\mathsf\{msPRL\}\}\\limits\. A pocket representation learning module is pretrained to encode protein pockets into multi\-scale atom\-level and residue\-level features\.b,๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsmodel training and inference\. Conditioned on these fixed pocket features,๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsgenerates 3D ligands for target pockets through a pocket\-conditioned diffusion model\.c,Generation\-time molecule optimization,๐๐บ๐ฎ๐ฏ๐ณ\\mathop\{\\mathsf\{paOPT\}\}\\limits\. During inference,๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsoptimizes the original trajectory in a training\-free manner to steer generated molecules toward improved ADMET properties\.
## Related Work
##### Generative Models for SBDD
Structure\-based drug design \(SBDD\) has leveraged various generative modeling approaches to design ligands tailored to specific protein pockets\. Peng*et al\.*\[peng22pocket2mol\]constructed an encoder\-predictor framework๐ฏ๐๐ผ๐๐พ๐๐ค๐ฌ๐๐
\\mathop\{\\mathsf\{Pocket2Mol\}\}\\limitsthat can generate ligands in an autoregressive way\. However,๐ฏ๐๐ผ๐๐พ๐๐ค๐ฌ๐๐
\\mathop\{\\mathsf\{Pocket2Mol\}\}\\limitsrelies on the autoregressive sampling process, which tends to violate geometric constraints and produce ligands with limited interactions with the pocket, resulting in low binding affinity\. More recently, target\-aware diffusion models\[guan2023targetdiff,guan2023decompdiff,zhoudecompopt,gu2024aligning,huang2024protein\]for ligand generation have largely addressed these limitations, offering better geometric consistency and binding affinities\. Guan*et al\.*\[guan2023targetdiff\]pioneered this direction by introducing an E\(3\)\-equivariant conditional diffusion model\[fuchs2020se,satorras2021n,ho2020ddpm,hoogeboom2021argmax\]๐ณ๐บ๐๐๐พ๐๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{TargetDiff\}\}\\limitsthat jointly generates atomic coordinates of ligands in 3D and types conditioned on protein pocket structures\. It achieves this through careful design of equivariant score networks and denoising processes that respect geometric symmetries, thereby establishing a strong baseline for diffusion\-based SBDD methods\.
To further improve conformational stability and molecular validity, Guan*et al\.*\[guan2023decompdiff\]proposed๐ฃ๐พ๐ผ๐๐๐๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{DecompDiff\}\}\\limits\. Inspired by medicinal chemistry practice of scaffold\-and\-arm design\[schneider1999scaffold\],๐ฃ๐พ๐ผ๐๐๐๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{DecompDiff\}\}\\limitsdecomposes ligand molecules into scaffolds and arms and learns separate diffusion priors for each\. This decomposed design facilitates more structured generation and enhances efficiency in exploring chemical space\. Building upon the idea of a decomposed prior, Zhou*et al\.*\[zhoudecompopt\]introduced๐ฃ๐พ๐ผ๐๐๐๐ฎ๐๐\\mathop\{\\mathsf\{DecompOpt\}\}\\limits, which decomposes pocket\-ligand interactions into local subpocket\-arm interactions and employs iterative optimization over the arms\. However,๐ณ๐บ๐๐๐พ๐๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{TargetDiff\}\}\\limits,๐ฃ๐พ๐ผ๐๐๐๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{DecompDiff\}\}\\limitsand๐ฃ๐พ๐ผ๐๐๐๐ฎ๐๐\\mathop\{\\mathsf\{DecompOpt\}\}\\limitsprimarily rely on geometric features of the pocket and ligand\-only priors without directly encoding protein\-ligand interaction signals into the diffusion dynamics\. To explicitly leverage such interactions, Huang*et al\.*\[huang2024protein\]proposed๐จ๐ฏ๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{IPDiff\}\}\\limits, which redesigns both forward and reverse processes to be informed by a learnable binding\-affinity signal\. This adaptation guides the generation toward ligands with improved binding properties\. Whereas๐จ๐ฏ๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{IPDiff\}\}\\limitssteers trajectories using a specific supervised affinity signal, Gu*et al\.*\[gu2024aligning\]introduced a more general framework๐ ๐ซ๐จ๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{ALIDiff\}\}\\limits, which controls generation by aligning the reverse process of a pretrained diffusion model with desired preferences via Diffusion\-DPO\[wallace2024diffusion\]adapted to molecular space\. Beyond diffusion\-based methods, recent work has also explored flow matching as an alternative generative paradigm for SBDD\. For example, Cremer*et al\.*\[cremer2026flowr\]introducedFLOWR, which combines continuous and categorical flow matching with equivariant optimal transport for ligand generation\.
Despite the different designs of current generative models, existing methods entangle pocket encoding with ligand denoising through a single end\-to\-end generation objective, which can blur interaction signals\. This gap motivates our design of decoupling the two: obtain a stable pretrained pocket summary first, then use it as a fixed context, enabling denoising to focus solely on pocket\-ligand interactions\. Moreover, prior works process pocket information only at the atom level, which may miss higher\-level structural patterns encoded by residues\. In contrast, our approach combines both atom\- and residue\-level pocket information to capture pocket characteristics and protein\-ligand interactions at multiple scales\. Furthermore, most prior works focus on improving binding affinity while ignoring ADMET profiles\. Our approach seeks to address this issue by explicitly incorporating ADMET\-aware optimization into the generation process\.
##### Alignment for Diffusion Models
There is a growing body of research on adapting pretrained generative diffusion models to specific downstream objectives\. These approaches can be broadly categorized into training\-based and training\-free methods, depending on whether the model parameters are updated\. In training\-based methods, Prabhudesai*et al\.*\[prabhudesai2023aligning\]proposed directly backpropagating through differentiable reward functions into diffusion model parameters\. While effective when such reward functions exist, this approach is prone to reward hacking in the absence of strong regularization\. Black*et al\.*\[blacktraining\]applied policy gradient methods such as PPO to finetune diffusion models\. Although more stable than direct backpropagation, reinforcement learning finetuning remains resource\- and time\-intensive, limiting its flexibility for diverse downstream tasks\. By contrast, training\-free methods operate in a plug\-and\-play fashion, requiring minimal or no parameter updates\. Song*et al\.*\[song2023loss\]introduced reward\-guided sampling, incorporating estimated reward gradients as guidance terms during generation\. Li*et al\.*\[li2024derivative\]proposed a value\-based decoding strategy that selects denoised samples according to estimated value functions\. More recently, Ma*et al\.*\[ma2025inference\]framed alignment as a search problem and introduced a search\-over\-paths algorithm to exploit generation\-time scaling behaviors\. While these methods advance training\-free alignment, they share a key limitation: it is difficult to reliably assess the quality of generated samples before the generative process is complete\.
To address the above\-mentioned limitations, a recent emerging line of work formulates the entire generation process as an optimization problem, where the objective is to directly optimize task\-specific downstream performance of the final generated sample\. Concretely, in standard diffusion sampling, each denoising step draws the next state from the model\-predicted posterior by combining a deterministic update \(the posterior mean\) with stochastic noise injected by the sampler\. By treating the injected noise variables as optimization variables, one can perturb the denoising trajectory and thereby steer the final sample toward preferred outcomes while keeping the diffusion model fixed\. Eyring*et al\.*\[eyring2024reno\]optimized the noise variable at the initial step of diffusion generation against external reward functions, achieving improved alignment at inference\. Tang*et al\.*\[tanginference\]extended this by optimizing injected noise across the entire denoising trajectory, propagating reward signals on the final outputs back to noise inputs while keeping the model parameters fixed\. While technically sound and effective in the field of image generation, these methods could not be readily applied for generating new chemical entities \(NCEs\)\. In particular, they are typically designed for continuous Gaussian noise and differentiable objectives, making it nontrivial to handle discrete \(categorical\) distribution over atom types and to incorporate molecular property evaluators, which are often black\-box and non\-differentiable\. Our proposed optimization scheme builds upon the idea of noise optimization with several key novelties\. First, whereas prior methods operate only on continuous latent noise, we extend the optimization framework to handle the categorical sampling process for discrete atom types\. Second, while existing works mainly consider differentiable reward functions, we investigate practical solutions for optimization under non\-differentiable, black\-box evaluators so that the framework can be applied to real\-world scenarios\. Finally, to the best of our knowledge, this work is the first to explore noise optimization in the context of drug design, and we demonstrate that this can be an effective mechanism for aligning diffusion\-based generation of NCEs with specific objectives \(e\.g\., higher absorption and lower toxicity\)\.
## Materials
### Datasets
We use the training and testing sets of the widely\-adopted CrossDocked2020\[Francoeur2020\]benchmark, referred to as๐ข๐ฃ\\mathop\{\\mathsf\{CD\}\}\\limits, for model training and testing, respectively\. In addition, we also use a new, manually annotated dataset,๐ข๐ฃ๐ง\\mathop\{\\mathsf\{CDH\}\}\\limits, for testing\.
#### Training data
We follow the data preparation and splitting process of Luo*et al\.*\[luo2021sbdd\], refining 22\.5 million docked binding complexes to high\-quality docking poses\. For each protein in๐ข๐ฃ\\mathop\{\\mathsf\{CD\}\}\\limits, there are multiple ligand\-binding complexes:\(1\)determined experimentally and reported in the Protein Data Bank \(PDB\)\[berman2000protein\], and\(2\)obtained through simulation by redocking ligands into non\-cognate receptors\. We retain high\-quality poses, including experimentally determined and docked poses whose heavy\-atom RMSD < 1 ร
relative to the experimentally\-observed ligands pose from the PDB, and select diverse proteins with pairwise sequence identity below 30%\. After filtering, the dataset consists of 100,000 protein\-ligand pairs for training\. As๐ข๐ฃ\\mathop\{\\mathsf\{CD\}\}\\limitsdoes not provide explicit pocket regions, following Guan*et al\.*\[guan2023targetdiff\], we define the binding pocket as all residues and their atoms within a 10 ร
radius of its ligand\.
#### Testing data
##### ๐ข๐ฃ\\mathop\{\\mathsf\{CD\}\}\\limitstesting data
The๐ข๐ฃ\\mathop\{\\mathsf\{CD\}\}\\limitstest set is consistent with the data preparation and splitting process described in๐ข๐ฃ\\mathop\{\\mathsf\{CD\}\}\\limitsโs training data\. We use all 100 protein\-ligand complexes from๐ข๐ฃ\\mathop\{\\mathsf\{CD\}\}\\limitsโs test set for testing\. We refer to the known ligand provided in the test set as the reference ligand\. The latter is used only to identify the pocket region in the protein structure\. Once the pocket is identified, it is provided to๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsand๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsas input, while the reference ligand itself is excluded from the generation process\. The test set includes human proteins associated with oncological, neurodegenerative, infectious, cardiovascular, and metabolic diseases, as well as non\-human proteins such as bacterial, fungal, parasitic, viral, and plant proteins\.
##### ๐ข๐ฃ๐ง\\mathop\{\\mathsf\{CDH\}\}\\limitstesting data
While๐ข๐ฃ\\mathop\{\\mathsf\{CD\}\}\\limitsโs test set has been widely used in the research community, we discovered that it contains protein targets for non\-human species, such as plants and bacteria, which limits its utility for human therapeutics discovery\. To address this limitation, we manually curated a new benchmarking dataset based on๐ข๐ฃ\\mathop\{\\mathsf\{CD\}\}\\limits, referred to as๐ข๐ฃ๐ง\\mathop\{\\mathsf\{CDH\}\}\\limits, which contains only human targets\.๐ข๐ฃ๐ง\\mathop\{\\mathsf\{CDH\}\}\\limitsretains all 33 human targets from๐ข๐ฃ\\mathop\{\\mathsf\{CD\}\}\\limitsโs test set, and includes an additional 39 human targets \(9 orthologs, 30 new targets\) from the PDB\. These additional PDB targets are carefully selected to cover high\-burden diseases such as cancers, neurological disorders, and autoimmune and inflammatory diseases, making๐ข๐ฃ๐ง\\mathop\{\\mathsf\{CDH\}\}\\limitsmore relevant to human diseases\. Table[E](https://arxiv.org/html/2607.12349#A5)in Appendix[E](https://arxiv.org/html/2607.12349#A5)presents the๐ข๐ฃ๐ง\\mathop\{\\mathsf\{CDH\}\}\\limitsinformation\.
For all the targets in๐ข๐ฃ๐ง\\mathop\{\\mathsf\{CDH\}\}\\limits, we consider their relevant key ADMET properties, including:\(1\)carcinogenicity \(Carci\), the propensity of a molecule to cause cancer,\(2\)Ames mutagenicity \(Ames\), the property of a molecule likely to mutate DNA,\(3\)hERG inhibition \(hERG\), the potential cardiac side effects due to inhibition of the cardiac hERG \(KCNH2\) potassium channel,\(4\)human intestinal absorption \(HIA\), the ability of an oral drug to be absorbed through the intestinal lining, and\(5\)blood\-brain barrier permeability \(BBBP\), the ability to cross the blood\-brain barrier and enter the central nervous system\. Assessment of carcinogenicity, Ames mutagenicity, and hERG inhibition is essential to ensure drug safety by mitigating risks of tumorigenicity, genotoxicity, and cardiotoxicity\. Evaluation of human intestinal absorption and blood\-brain barrier permeability is critical to establish pharmacokinetic suitability and therapeutic accessibility of drug candidates\. These ADMET properties represent key determinants in drug discovery, guiding the identification and optimization of viable therapeutic candidates\. Unfortunately, current generative SBDD models do not proactively evaluate or optimize these properties, largely missing opportunities to increase the overall success rate of the drug discovery and development process\.๐ข๐ฃ๐ง\\mathop\{\\mathsf\{CDH\}\}\\limitsis constructed to enable such ADMET evaluation\.
### Baselines
To evaluate the effectiveness of๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsand๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsin generating ligands that bind to target protein pockets, we compare them against the following state\-of\-the\-art baselines for SBDD:๐ ๐ฑ\\mathop\{\\mathsf\{AR\}\}\\limits\[luo2021sbdd\],๐ฏ๐๐ผ๐๐พ๐๐ค๐ฌ๐๐
\\mathop\{\\mathsf\{Pocket2Mol\}\}\\limits\[peng22pocket2mol\],๐ฃ๐๐ฟ๐ฟ๐ฒ๐ก๐ฃ๐ฃ\\mathop\{\\mathsf\{DiffSBDD\}\}\\limits\[schneuing2022structure\],๐ณ๐บ๐๐๐พ๐๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{TargetDiff\}\}\\limits\[guan2023targetdiff\],๐ฃ๐พ๐ผ๐๐๐๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{DecompDiff\}\}\\limits\[guan2023decompdiff\],๐ฃ๐พ๐ผ๐๐๐๐ฎ๐๐\\mathop\{\\mathsf\{DecompOpt\}\}\\limits\[zhoudecompopt\],๐จ๐ฏ๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{IPDiff\}\}\\limits\[huang2024protein\]and๐ ๐ซ๐จ๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{ALIDiff\}\}\\limits\[gu2024aligning\], where๐ ๐ฑ\\mathop\{\\mathsf\{AR\}\}\\limitsand๐ฏ๐๐ผ๐๐พ๐๐ค๐ฌ๐๐
\\mathop\{\\mathsf\{Pocket2Mol\}\}\\limitsare non\-diffusion methods, and the rest are all diffusion\-based methods\. These methods generate 3D binding ligands conditioned on the pockets of protein targets\. We choose these baselines because they are well\-established SBDD methods with strong performance on๐ข๐ฃ\\mathop\{\\mathsf\{CD\}\}\\limits\. We used their author\-provided implementations with released checkpoints and default inference settings for all baselines to conduct computational experiments on๐ข๐ฃ๐ง\\mathop\{\\mathsf\{CDH\}\}\\limits\(except๐ ๐ซ๐จ๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{ALIDiff\}\}\\limits, which does not provide checkpoints\)\.
### Model Training and Evaluation
All the models, including๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limits,๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits, and the baselines, are trained over๐ข๐ฃ\\mathop\{\\mathsf\{CD\}\}\\limitstraining data, and tested on the๐ข๐ฃ\\mathop\{\\mathsf\{CD\}\}\\limitstest set and๐ข๐ฃ๐ง\\mathop\{\\mathsf\{CDH\}\}\\limits\. In๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limits,๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limitsand๐๐ผ๐ซ๐ฆ\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{pcLG\}\}\\limits$\}\}\\limitshave different training objectives, thus,๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limitsis trained using only the protein pockets from๐ข๐ฃ\\mathop\{\\mathsf\{CD\}\}\\limitstraining set, and๐๐ผ๐ซ๐ฆ\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{pcLG\}\}\\limits$\}\}\\limitsis trained using the protein\-ligand complexes\. In๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits, we focus on the targets in๐ข๐ฃ๐ง\\mathop\{\\mathsf\{CDH\}\}\\limitsand their ADMET property optimization, and its generation\-time optimization does not need model finetuning\. For each target in๐ข๐ฃ๐ง\\mathop\{\\mathsf\{CDH\}\}\\limits,๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsoptimizes one particular ADMET property key to that target during the generation time\. The properties to optimize for different disease targets are presented in Table[E](https://arxiv.org/html/2607.12349#A5)in Appendix[E](https://arxiv.org/html/2607.12349#A5)\. Thus, the evaluation is organized with respect to each ADMET property across multiple relevant targets\.
### Evaluation Metrics
#### General ligand properties
To predict binding affinity, following the literature\[guan2023targetdiff,guan2023decompdiff\], we use Vina Scores \(Vina S\) calculated by AutoDock Vina\[Eberhardt2021\], which evaluates the quality of the binding poses of generated ligands against protein targets\. In addition, as suggested in the literature\[guan2023targetdiff,guan2023decompdiff\], we optimize the poses of the generated 3D ligands using a local energy minimization algorithm and a docking algorithm, as implemented in AutoDock Vina\[Eberhardt2021\]\. We then evaluate the predicted binding affinities of these optimized poses via two metrics: Vina Minimization \(Vina M\) based on local energy minimization, and Vina Dock \(Vina D\) based on docking\. Lower Vina scores indicate stronger binding affinities\. Based on Vina D, following the literature\[guan2023targetdiff,guan2023decompdiff\], we also measure the percentage of how many generated ligands across all the targets bind better than their reference ligand, referred to as High Affinity percentage \(HA%\)\. Higher HA% indicates a better capacity to generate ligands above the reference affinities\.
For drug\-likeness, we evaluate whether the generated ligands are drug\-like using the quantitative estimate of drug\-likeness \(QED\)\[Bickerton2012\]and synthesizable using a synthetic accessibility \(SA\) score\[Ertl2009\]\. We also calculate the diversity among generated ligands of each target, which is a per\-target measurement and defined as the average pairwise Tanimoto distances\[bajusz2015tanimoto\]derived using 2048\-bit fingerprints as implemented within RDKit\[rdkit\]\. Higher diversity indicates a better ability to explore broader chemical space\. To jointly consider predicted binding affinity, drug\-likeness, and synthetic feasibility, following previous work\[guan2023targetdiff,guan2023decompdiff\], we evaluate the success rate \(SR%\) calculated as the percentage of all generated ligands across all targets with Vina D<<\-8\.18, QED\>\>0\.25, and SA\>\>0\.59\. Higher SR% indicates a better ability to generate high\-quality ligands\. The Vina D score threshold \-8\.18, which corresponds to a binding affinity of less than 1ฮผ\\muM, is widely recognized in medicinal chemistry as an indicator of moderate biological activity\[zhoudecompopt\]\. The QED and SA score thresholds of 0\.25 and 0\.59 are set as the 10th percentile of approved drugs in DrugCentral\[ursu2016drugcentral\], enforcing reasonable lower bounds to maximize the consideration of potentially promising ligands\.
#### ADMET properties
We evaluate the generated ligands on five ADMET properties: Carci, Ames, hERG, HIA, and BBBP\. We useADMET\-AI\[swanson2024admet\]to estimate a probabilistic score for each property\. For Carci, Ames, and hERG, lower scores are preferred, indicating lower toxicity risks, while for HIA and BBBP, higher scores are more desired, indicating higher human intestinal absorption and blood\-brain barrier permeability, respectively\.
#### Ligand structures
We also evaluate the quality of generated ligands by analyzing their stability and 3D structural quality\. Specifically, to evaluate stability, we calculate both atomic level and molecular level stability, as described in literature\[hoogeboom22diff\]\. Atomic level stability is defined by the percentage of atoms that maintain correct valency, while molecular level stability measures the percentage of molecules in which all the atoms are stable, both for all the generated ligands and for all the targets\. We evaluate the 3D structures of generated ligands using the same metrics as in Peng*et al\.*\[peng2023moldiff\]\. We calculate the Jensen\-Shannon \(JS\) divergences for bond lengths, bond angles, and dihedral angles, which measure how far the distributions of these properties of generated ligands are compared with those of real molecules \(i\.e\., training molecules\) regarding the 3D structures\.
## Results
We assess the performance of the baselines and๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitson the 100 test pockets from๐ข๐ฃ\\mathop\{\\mathsf\{CD\}\}\\limitstest set and the 72 curated test pockets from๐ข๐ฃ๐ง\\mathop\{\\mathsf\{CDH\}\}\\limits, using the evaluation metrics with 100 ligands generated per pocket by these methods\. We evaluate๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsusing the๐ข๐ฃ๐ง\\mathop\{\\mathsf\{CDH\}\}\\limits, optimizing the 100 ligands per pocket with respect to the ADMET properties most relevant to its corresponding therapeutic indication\. Here, we present the results on๐ข๐ฃ๐ง\\mathop\{\\mathsf\{CDH\}\}\\limits; the results on๐ข๐ฃ\\mathop\{\\mathsf\{CD\}\}\\limitsare discussed in AppendixLABEL:supp:cd\.
### Overall Comparison on๐ข๐ฃ๐ง\\mathop\{\\mathsf\{CDH\}\}\\limits
Table 1:Comparison on๐ข๐ฃ๐ง\\mathop\{\\mathsf\{CDH\}\}\\limitsMethodVina Sโ\\downarrowVina Mโ\\downarrowVina Dโ\\downarrowHA%โ\\uparrowQEDโ\\uparrowSAโ\\uparrowDivโ\\uparrowSR%โ\\uparrowAvg\.Med\.Avg\.Med\.Avg\.Med\.Avg\.Med\.Avg\.Med\.Avg\.Med\.Avg\.Med\.Reference\-7\.45\-7\.35\-7\.89\-7\.63\-8\.01\-7\.71\-\-0\.460\.460\.730\.76\-\-29\.2๐ ๐ฑ\\mathop\{\\mathsf\{AR\}\}\\limits\-6\.63\-6\.51\-6\.95\-6\.69\-7\.43\-7\.2137\.021\.60\.520\.530\.630\.630\.690\.6912\.2๐ฏ๐๐ผ๐๐พ๐๐ค๐ฌ๐๐
\\mathop\{\\mathsf\{Pocket2Mol\}\}\\limits\-5\.28\-5\.02\-6\.44\-6\.16\-7\.23\-7\.0033\.115\.70\.610\.610\.790\.800\.790\.8129\.7๐ฃ๐๐ฟ๐ฟ๐ฒ๐ก๐ฃ๐ฃ\\mathop\{\\mathsf\{DiffSBDD\}\}\\limits\-3\.24\-5\.12\-5\.27\-5\.94\-7\.02\-7\.3034\.121\.90\.480\.490\.630\.610\.780\.7511\.3๐ณ๐บ๐๐๐พ๐๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{TargetDiff\}\}\\limits\-6\.38\-6\.67\-7\.18\-7\.25\-8\.24\-8\.1947\.143\.90\.460\.470\.580\.570\.710\.7012\.3๐ฃ๐พ๐ผ๐๐๐๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{DecompDiff\}\}\\limits\-6\.17\-6\.02\-6\.96\-6\.76\-7\.88\-7\.8349\.845\.50\.480\.480\.640\.630\.670\.6618\.3๐ฃ๐พ๐ผ๐๐๐๐ฎ๐๐\\mathop\{\\mathsf\{DecompOpt\}\}\\limits\-6\.34\-6\.22\-7\.04\-6\.88\-7\.98\-7\.8351\.752\.90\.480\.480\.650\.640\.670\.6720\.5๐จ๐ฏ๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{IPDiff\}\}\\limits\-8\.02\-8\.15\-8\.43\-8\.30\-9\.24\-8\.9162\.066\.30\.460\.460\.550\.540\.720\.7110\.0๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limits\-6\.80\-7\.10\-7\.69\-7\.78\-8\.85\-8\.7162\.274\.60\.460\.450\.590\.580\.600\.5822\.0
- โขโโColumns represent: โVina Sโ: the predicted binding affinities between the initially generated poses of ligands and the protein pockets; โVina Mโ: the predicted binding affinities between the poses after local structure minimization and the protein pockets; โVina Dโ: the predicted binding affinities between the poses determined by AutoDock Vina and the protein pockets; โHAโ: the percentage of generated ligands with Vina D lower than those of reference ligands; โQEDโ: the quantitative estimate of drug\-likeness; โSAโ: the synthetic accessibility score; โDivโ: the diversity among generated ligands; โSRโ: the percentage of generated molecules with predicted binding affinities, QED and SA above certain thresholds\. The best values are inbold, and the second\-best values areunderlined\. Rows ingrayare excluded from the comparison\.โ\\uparrow/โ\\downarrowindicate that higher / lower values are better\.
Table[1](https://arxiv.org/html/2607.12349#Sx4.T1)presents a comprehensive comparison between๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsand all the baselines in terms of generating drug\-like and diverse ligands that effectively bind to protein pockets of๐ข๐ฃ๐ง\\mathop\{\\mathsf\{CDH\}\}\\limits\.๐จ๐ฏ๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{IPDiff\}\}\\limitsis reported for completeness but excluded from best/second\-best marking because ligands generated by this method exhibit simplified structures with excessive carbons and single bonds \(Details are in the Analysis of Atom and Bond Composition on๐ข๐ฃ๐ง\\mathop\{\\mathsf\{CDH\}\}\\limitssection\)\. In terms of Vina S \(Avg/Med\),๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsoutperforms all the baseline methods, and also achieves the highest HA%, with a 2\.6% and 24\.9% improvement over the best baselines on Vina S and HA%, respectively\. This suggests that๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsbetter guides ligand atoms toward regions where they can engage in interactions with the pocket\. After energy minimization and re\-docking,๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsremains superior in terms of Vina M and Vina D scores, with 7\.1% and 7\.4% improvement over the best baselines, respectively\. This indicates the generated initial poses offer strong starting points for further refinement\.๐ ๐ฑ\\mathop\{\\mathsf\{AR\}\}\\limitsis the second\-best model in terms of Vina S\.๐ ๐ฑ\\mathop\{\\mathsf\{AR\}\}\\limitslearns pocket\-anchored atom type density maps, making it easier to place atoms and achieve high binding affinities\. In terms of Vina M and Vina D,๐ณ๐บ๐๐๐พ๐๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{TargetDiff\}\}\\limitsachieves the second\-best performance\. Unlike๐ ๐ฑ\\mathop\{\\mathsf\{AR\}\}\\limits, whose auto\-regressive sampling strategy suffers from high variance in pose quality,๐ณ๐บ๐๐๐พ๐๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{TargetDiff\}\}\\limitsgenerates more consistent poses, improving Vina M and Vina D over๐ ๐ฑ\\mathop\{\\mathsf\{AR\}\}\\limitsbut still underperforming๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limits\. Among the baseline methods,๐ฃ๐พ๐ผ๐๐๐๐ฎ๐๐\\mathop\{\\mathsf\{DecompOpt\}\}\\limitsspecifically performs iterative, docking\-guided arm optimization, aiming to achieve high binding affinities and thus high Vina scores\. Instead,๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsdoes not explicitly conduct such optimization โ the fact that it still outperforms๐ฃ๐พ๐ผ๐๐๐๐ฎ๐๐\\mathop\{\\mathsf\{DecompOpt\}\}\\limitsin binding affinity scores indicates๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsencodes pocket information and pocket\-ligand interactions better than baselines\. Compared with diffusion models that do not perform any Vina\-based optimization \(๐ฃ๐๐ฟ๐ฟ๐ฒ๐ก๐ฃ๐ฃ\\mathop\{\\mathsf\{DiffSBDD\}\}\\limits,๐ณ๐บ๐๐๐พ๐๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{TargetDiff\}\}\\limits,๐ฃ๐พ๐ผ๐๐๐๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{DecompDiff\}\}\\limits\),๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsattains the best binding affinity scores while being more sample efficient\. The key difference is that๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsconditions generation on a pretrained pocket embedding, so pocket information is available*a priori*and does not need to be encoded every timestep during sampling, reducing overhead and improving pose quality\. In addition,๐ฃ๐๐ฟ๐ฟ๐ฒ๐ก๐ฃ๐ฃ\\mathop\{\\mathsf\{DiffSBDD\}\}\\limitsshows substantially lower Vina S than other methods\. A possible reason is coarse encoding of pocket structure, making๐ฃ๐๐ฟ๐ฟ๐ฒ๐ก๐ฃ๐ฃ\\mathop\{\\mathsf\{DiffSBDD\}\}\\limitsignore some important structural information and potential interaction sites\. This leads๐ฃ๐๐ฟ๐ฟ๐ฒ๐ก๐ฃ๐ฃ\\mathop\{\\mathsf\{DiffSBDD\}\}\\limitsto struggle to place ligands correctly within the pocket, causing steric clashes, suboptimal functional\-group orientations, and consequently higher Vina scores\. In contrast,๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsavoids this issue through its dedicated pocket encoder\.
The๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsmethod achieves QED and SA scores comparable with those of other diffusion baselines\. Notably, the two autoregressive models \(๐ ๐ฑ\\mathop\{\\mathsf\{AR\}\}\\limitsand๐ฏ๐๐ผ๐๐พ๐๐ค๐ฌ๐๐
\\mathop\{\\mathsf\{Pocket2Mol\}\}\\limits\) exhibit QED and SA scores significantly higher than those of all diffusion\-based models\. These two methods tend to generate smaller and simpler molecules via an atom\-by\-atom autoregression, with an average heavy\-atom count of 17 and 20, respectively, compared to 26 from other diffusion\-based baselines and 28 from the references\. While such molecules are favored by QED and SA measurement due to small size, simple structures and limited scaffolds\[bickerton2012quantifying,ertl2009estimation\],๐ ๐ฑ\\mathop\{\\mathsf\{AR\}\}\\limitsand๐ฏ๐๐ผ๐๐พ๐๐ค๐ฌ๐๐
\\mathop\{\\mathsf\{Pocket2Mol\}\}\\limitsdo not generally enjoy the same favor in terms of binding affinities\. Among the diffusion\-based methods, the decomposition baselines \(๐ฃ๐พ๐ผ๐๐๐๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{DecompDiff\}\}\\limitsand๐ฃ๐พ๐ผ๐๐๐๐ฎ๐๐\\mathop\{\\mathsf\{DecompOpt\}\}\\limits\) achieve highest QED and SA scores\. Their decomposed, multiple priors for different atom roles bias the generation of molecules toward those with more drug\-like, synthesizable fragments, yielding higher QED and SA scores than those from other diffusion models of a single prior\.๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsemphasizes binding\-quality during generation, and therefore, shows moderate QED and SA with high binding affinities\.
We observe that๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsโโs diversity slightly falls behind the baselines\. However,๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsโโs diversity is still acceptable given its strengths in generating high\-affinity molecules, and it is comparable with those of๐ฃ๐พ๐ผ๐๐๐๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{DecompDiff\}\}\\limitsand๐ฃ๐พ๐ผ๐๐๐๐ฎ๐๐\\mathop\{\\mathsf\{DecompOpt\}\}\\limits\. We hypothesize that the diversity is limited because the pocket is pre\-encoded as a fixed condition for diffusion\. A static pocket representation may sharpen the denoising landscape, limiting the chemical space that๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsmay visit\. In contrast, the baselines that jointly diffuse ligand and pocket allow the pocket representation to adapt to noisy states of ligands, and thus, achieve higher ligand diversity\. In terms of SR%,๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsattains the second\-best SR%, notably above all diffusion baselines, indicating that it balances design objectives of drugs effectively and yields realistic ligand candidates that are both effective binders and chemically viable\.๐ฏ๐๐ผ๐๐พ๐๐ค๐ฌ๐๐
\\mathop\{\\mathsf\{Pocket2Mol\}\}\\limitsachieves the highest SR%, consistent with its strong QED and SA performance\. However,๐ฏ๐๐ผ๐๐พ๐๐ค๐ฌ๐๐
\\mathop\{\\mathsf\{Pocket2Mol\}\}\\limitssignificantly underperforms๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsin terms of Vina scores\.
### Overall Comparison after ADMET Optimization
##### Overall Performance of๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits
Table 2:Comparison on๐ข๐ฃ๐ง\\mathop\{\\mathsf\{CDH\}\}\\limitswith optimized property CarciMethodVina Sโ\\downarrowVina Mโ\\downarrowVina Dโ\\downarrowHA%โ\\uparrowQEDโ\\uparrowSAโ\\uparrowDivโ\\uparrowSR%โ\\uparrowCarciโ\\downarrowAvg\.Med\.Avg\.Med\.Avg\.Med\.Avg\.Med\.Avg\.Med\.Avg\.Med\.Avg\.Med\.Avg\.Med\.Reference\-7\.53\-7\.45\-7\.93\-7\.68\-8\.00\-7\.71\-\-0\.450\.430\.730\.77\-\-25\.00\.190\.14๐ ๐ฑ\\mathop\{\\mathsf\{AR\}\}\\limits\-6\.52\-6\.44\-6\.87\-6\.61\-7\.37\-7\.0936\.721\.00\.520\.530\.630\.630\.680\.6811\.20\.140\.07๐ฏ๐๐ผ๐๐พ๐๐ค๐ฌ๐๐
\\mathop\{\\mathsf\{Pocket2Mol\}\}\\limits\-5\.30\-5\.04\-6\.46\-6\.15\-7\.25\-6\.9634\.514\.50\.610\.620\.790\.800\.790\.8029\.50\.180\.13๐ฃ๐๐ฟ๐ฟ๐ฒ๐ก๐ฃ๐ฃ\\mathop\{\\mathsf\{DiffSBDD\}\}\\limits\-3\.11\-5\.12\-5\.22\-5\.92\-6\.88\-7\.2534\.320\.30\.490\.500\.630\.620\.780\.7611\.60\.180\.12๐ณ๐บ๐๐๐พ๐๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{TargetDiff\}\}\\limits\-6\.37\-6\.60\-7\.17\-7\.23\-8\.20\-8\.1547\.941\.10\.460\.470\.590\.580\.710\.6912\.30\.200\.15๐ฃ๐พ๐ผ๐๐๐๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{DecompDiff\}\}\\limits\-6\.12\-5\.88\-6\.93\-6\.63\-7\.86\-7\.7448\.549\.50\.480\.490\.640\.620\.660\.6617\.10\.200\.15๐ฃ๐พ๐ผ๐๐๐๐ฎ๐๐\\mathop\{\\mathsf\{DecompOpt\}\}\\limits\-6\.30\-6\.14\-7\.02\-6\.79\-8\.00\-7\.7854\.156\.10\.480\.490\.640\.640\.670\.6619\.10\.200\.15๐จ๐ฏ๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{IPDiff\}\}\\limits\-7\.98\-8\.23\-8\.41\-8\.40\-9\.22\-8\.9162\.166\.30\.460\.470\.550\.540\.720\.7110\.20\.220\.18๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limits\-6\.66\-6\.80\-7\.46\-7\.45\-8\.48\-8\.4255\.962\.30\.490\.510\.620\.600\.650\.6223\.00\.220\.22๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-6\.55\-6\.80\-7\.35\-7\.38\-8\.37\-8\.3654\.256\.60\.520\.530\.600\.580\.630\.6118\.20\.060\.06- โขโโColumns represent: โVina Sโ: the predicted binding affinities between the initially generated poses of ligands and the protein pockets; โVina Mโ: the predicted binding affinities between the poses after local structure minimization and the protein pockets; โVina Dโ: the predicted binding affinities between the poses determined by AutoDock Vina and the protein pockets; โHAโ: the percentage of generated ligands with Vina D lower than those of reference ligands; โQEDโ: the quantitative estimate of drug\-likeness; โSAโ: the synthetic accessibility score; โDivโ: the diversity among generated ligands; โSRโ: the percentage of generated molecules with binding affinities, QED, and SA above a certain threshold\. โCarciโ: the predicted carcinogenicity score for generated molecules; The best values are inbold, and the second\-best values areunderlined\. Rows ingrayare excluded from the comparison\.โ\\uparrow/โ\\downarrowindicate that higher / lower values are better\.
The๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsmethod maintains strong performance on general molecular metrics compared with all evaluated baselines, ranking among the top methods in predicted binding affinity \(Vina S/M/D\), drug\-likeness \(QED\), and synthetic accessibility \(SA\)\. To examine๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsโs performance in more detail, in Table[2](https://arxiv.org/html/2607.12349#Sx4.T2), we take the carcinogenicity optimization setting as an example, and compare the general metrics of๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsligands with those of๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsligands to ensure that the ADMET optimization procedure does not disrupt the desirable structures of initial๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsligands\. In general, the results show that ligands produced by๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsachieve Vina S/M/D, QED, and SA scores that are on par with, and in some cases exceed, those of๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limits\. As shown in Table[2](https://arxiv.org/html/2607.12349#Sx4.T2),๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsreduces the average Carci score from 0\.22 to 0\.06 \(by 73%\), which is the best among all methods, while the average Vina S/M/D scores change only slightly by less than 2\.5%\. Despite this marginal degradation, the resulting Vina scores of๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsremain stronger than all other baselines \(except๐จ๐ฏ๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{IPDiff\}\}\\limits\)\. The๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsand๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsmethods perform very similarly across QED and SA scores in Tables[2](https://arxiv.org/html/2607.12349#Sx4.T2), indicating that drug\-likeness and synthetic accessibility are well preserved after optimization\. This preservation can be attributed to the noise optimization design in๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits: rather than directly modifying the molecular structures of ligands generated by๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limits, the optimization is performed in the diffusion noise space, refining the generation process while keeping ligands within๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsโs learned generative distribution\. We note that the diversity of๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsligands remains similar to that of๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsligands, suggesting that๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsenhances ADMET properties without collapsing the generated molecules into a narrow structural family\. We also summarize the results of optimizing other ADMET properties \(Ames, hERG, BBBP and HIA\) in TablesLABEL:tbl:overall\_results\_ames\-LABEL:tbl:overall\_results\_hiain AppendixLABEL:supp:opt\_results, and observe a similar pattern to carcinogenicity optimization:๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsconsistently improves the targeted ADMET score substantially while only modestly affecting binding affinity scores and other general molecular metrics\.
##### ADMET Property Optimization
Our results demonstrate that, leveraging noise optimization, the sampling process of๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsis guided towards a chemical subspace with more favorable ADMET properties, facilitating efficient discovery of safer and more effective drug candidates\. Table[3](https://arxiv.org/html/2607.12349#Sx4.T3)presents a comparison of generated molecules across different ADMET properties of interest, where each property score is optimized in a separate experiment\. The reported results of๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsare obtained by evaluating the best\-performing outcomes across multiple optimization iterations\. As the results show,๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsconsistently generates molecules with substantially more favorable ADMET properties compared to those from the baselines\. Specifically,๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsproduces ligands with the lowest predicted carcinogenicity and Ames mutagenicity, achieving roughly 50% to 60% better average Carci and Ames scores, compared to all evaluated baselines\. This indicates๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsgenerates drug candidates with much lower predicted risks of cancer induction and DNA damage across evaluated protein targets\. Similarly, Table[3](https://arxiv.org/html/2607.12349#Sx4.T3)demonstrates that๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsdecreases the hERG property score by 32%, relative to the next best baseline \(๐ฏ๐๐ผ๐๐พ๐๐ค๐ฌ๐๐
\\mathop\{\\mathsf\{Pocket2Mol\}\}\\limits\), indicating considerable reduction in the predicted cardiotoxicity of the generated ligands\. Finally,๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsenhances absorption\-related properties, achieving notable improvements in both BBBP and HIA, compared to other baseline methods\. Taken together, these results highlight the effectiveness of๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsas an optimization method targeting desired ADMET objectives\.
Table 3:Comparison of ADMET property scores on๐ข๐ฃ๐ง\\mathop\{\\mathsf\{CDH\}\}\\limitsMethodCarciโ\\downarrowAmesโ\\downarrowhERGโ\\downarrowBBBPโ\\uparrowHIAโ\\uparrowAvg\.Med\.Avg\.Med\.Avg\.Med\.Avg\.Med\.Avg\.Med\.Reference0\.190\.140\.330\.280\.440\.380\.660\.700\.730\.98๐ ๐ฑ\\mathop\{\\mathsf\{AR\}\}\\limits0\.140\.070\.480\.470\.310\.210\.700\.770\.880\.99๐ฏ๐๐ผ๐๐พ๐๐ค๐ฌ๐๐
\\mathop\{\\mathsf\{Pocket2Mol\}\}\\limits0\.180\.130\.420\.400\.220\.110\.760\.840\.941\.00๐ฃ๐๐ฟ๐ฟ๐ฒ๐ก๐ฃ๐ฃ\\mathop\{\\mathsf\{DiffSBDD\}\}\\limits0\.180\.120\.460\.430\.300\.210\.710\.770\.911\.00๐ณ๐บ๐๐๐พ๐๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{TargetDiff\}\}\\limits0\.200\.150\.360\.290\.350\.260\.680\.720\.921\.00๐ฃ๐พ๐ผ๐๐๐๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{DecompDiff\}\}\\limits0\.200\.150\.340\.270\.380\.280\.700\.750\.820\.99๐ฃ๐พ๐ผ๐๐๐๐ฎ๐๐\\mathop\{\\mathsf\{DecompOpt\}\}\\limits0\.200\.150\.350\.290\.370\.260\.690\.750\.821\.00๐จ๐ฏ๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{IPDiff\}\}\\limits0\.220\.180\.240\.160\.430\.460\.800\.860\.971\.00๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limits0\.220\.220\.370\.380\.430\.440\.640\.620\.930\.98๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits0\.060\.060\.100\.090\.150\.120\.890\.900\.991\.00- โขColumns represent: โCarciโ: the carcinogenicity; โAmesโ: Ames mutagenicity; โhERGโ: hERG inhibition; โBBBPโ: blood\-brain barrier permeability; โHIAโ: human intestinal absorption\. The best values are inbold, and the second\-best values areunderlined\.โ\\uparrow/โ\\downarrowindicate that higher / lower values are better\.
##### Change in ADMET Properties over Optimization Iterations
To gain deeper insight into how the optimized as well as other non\-optimized ADMET properties are affected during the optimization process, in Figure[2](https://arxiv.org/html/2607.12349#Sx4.F2), we visualize the changes in ADMET property scores of๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsgenerated ligands across optimization iterations\. Each property in the corresponding subplot is optimized independently, while the trajectories of non\-optimized properties are tracked in parallel\. For optimized properties, the scores consistently shift in the intended direction of improvement as the number of iterations increases โ carcinogenicity, Ames, and hERG scores gradually decline, while HIA and BBBP scores steadily increase\. In contrast, the trajectories of non\-optimized properties remain largely stable, indicating that, in๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits, optimizing one property does not substantially compromise other properties\. Nevertheless, we observe mild tradeoffs across certain properties\. For instance, reducing hERG tends to slightly lower BBBP and HIA scores, and conversely, improving absorption\-related properties \(HIA, BBBP\) can increase hERG\. This coupling reflects common physicochemical and structural factors underlying these ADMET properties\. A moleculeโs ability to permeate intestinal and blood\-brain barriers is often influenced by its lipophilicity, polarity, and molecular size\[weiss2024balanced\]\. However, these same factors also impact a moleculeโs ability to bind the hERG potassium channel\. This highlights a key challenge in lead optimization: each structural modification has myriad effects that must be carefully monitored and controlled\. Disentangling such effects remains a difficult problem, which we leave for future exploration, possibly through multi\-objective optimization\.

Fig\. 2:Changes in ADMET property scores over๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsoptimization iterations\. In each subfigure, the solid line indicates the property being optimized, while the dashed lines indicate the other properties not targeted for optimization\.๐\\mathbf\{a\}\-๐\\mathbf\{e\}show the trends when optimizing Carci, Ames, hERG, HIA, and BBBP, respectively\.
### Comparison on Molecular Structures of Generated Ligands
Table 4:Comparison on molecular structures of generated ligands on๐ข๐ฃ๐ง\\mathop\{\\mathsf\{CDH\}\}\\limitsGroupMetric๐ ๐ฑ\\mathop\{\\mathsf\{AR\}\}\\limits๐ฏ๐๐ผ๐๐พ๐๐ค๐ฌ๐๐
\\mathop\{\\mathsf\{Pocket2Mol\}\}\\limits๐ฃ๐๐ฟ๐ฟ๐ฒ๐ก๐ฃ๐ฃ\\mathop\{\\mathsf\{DiffSBDD\}\}\\limits๐ณ๐บ๐๐๐พ๐๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{TargetDiff\}\}\\limits๐ฃ๐พ๐ผ๐๐๐๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{DecompDiff\}\}\\limits๐จ๐ฏ๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{IPDiff\}\}\\limits๐ฃ๐พ๐ผ๐๐๐๐ฎ๐๐\\mathop\{\\mathsf\{DecompOpt\}\}\\limits๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limits๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsStabilityAtom stability \(โ\\uparrow\)0\.9220\.8710\.8890\.9490\.9060\.9390\.9160\.9350\.937Molecule stability \(โ\\uparrow\)0\.4490\.1710\.2200\.3890\.2350\.3560\.2670\.3300\.3493D structuresJS\. bond lengths \(โ\\downarrow\)0\.4630\.4370\.3290\.3360\.2980\.4810\.2850\.2720\.272JS\. bond angles \(โ\\downarrow\)0\.3720\.2440\.2980\.2340\.1640\.3200\.1590\.1870\.190JS\. dihedral angles \(โ\\downarrow\)0\.4080\.2330\.2900\.2580\.2020\.3250\.1990\.1460\.147
- Rows represent: โatom stabilityโ: the proportion of stable atoms that have the correct valency; โmolecule stabilityโ: the proportion of generated ligands with all atoms stable; โJS\. bond lengths/bond angles/dihedral anglesโ: the JensenโShannon \(JS\) divergences of bond lengths, bond angles, and dihedral angles between generated ligands and training ligands\. The best values are inbold, and the second\-best values areunderlined\.โ\\uparrow/โ\\downarrowindicate that higher / lower values are better\.
Table[4](https://arxiv.org/html/2607.12349#Sx4.T4)presents stability and 3D structure metrics for generated ligands, evaluating structural plausibility, such as proper atomic valences and realistic bond angles and dihedral angles geometry\. Notably,๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsand๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsachieve atom\-stability scores of 0\.935 and 0\.937, respectively, indicating strong adherence to correct valency at the atom level\. The๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsand๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsmethods also deliver competitive molecule stability, suggesting their ligands are chemically plausible\. Interestingly,๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsintroduces additional gains over๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitson both atom and molecule stability, suggesting the guidance from ADMET properties also biases generation towards more realistic structures with fewer valence violations and lower bond strain\.
Table[4](https://arxiv.org/html/2607.12349#Sx4.T4)also presents Jensen\-Shannon \(JS\) divergences\[lin2002divergence\]between the generated ligands and the training data ligands for bond lengths, bond angles, and dihedral angles\. The๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsand๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsmethods generate ligands with the lowest divergences from the training ligand distribution for bond lengths and dihedral angles, indicating both methods can generate molecules with realistic chemical structures\. Furthermore, the bond angle distributions for๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsand๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsare only slightly more divergent from the training distribution than the best baseline,๐ฃ๐พ๐ผ๐๐๐๐ฎ๐๐\\mathop\{\\mathsf\{DecompOpt\}\}\\limits\. Overall,๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsand๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsdemonstrate a strong ability to generate ligands with high\-quality chemical structures, mimicking the real\-world structural qualities of the training ligands\.
Fig\. 3:Comparison on filter pass rates of generated ligands to marketed drugs from ChemBL\[gaulton2012chembl\]\. Rule of Five, Rule of Ghose, Rule of Veber, and Rule of ZINC are common drug\-likeness rules, covering physicochemical properties, lipophilicity, hydrogen\-bonding capacity, molecular size, flexibility, and bioavailability\-related characteristics; BMS, PAINS, SureChEMBL, and NIBR are pharmaceutical structural alerts, covering undesirable, reactive, promiscuous, or assay\-interfering substructures; Complexity, Bredt, and Molecular Graph are cheminformatics validity tests, covering molecular complexity, strained structural motifs, and chemically implausible or unstable molecular graphs\.Figure[3](https://arxiv.org/html/2607.12349#Sx4.F3)presents the performance of generated ligands from๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsand๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitson several drug discovery filters, applying common drug\-likeness rules \(Rule of Five\[lipinski2004lead\], Rule of Ghose\[ghose1999knowledge\], Rule of Veber\[veber2002molecular\], Rule of ZINC\[irwin2005zinc,sterling2015zinc,irwin2020zinc20\]\), pharmaceutical structural alerts \(BMS\[pearce2006empirical\], PAINS\[baell2010new\], SureChEMBL\[papadatos2016surechembl\], NIBR\[schuffenhauer2020evolution\]\), and cheminformatics validity tests \(Complexity\[bertz1981first\], Bredt\[kobrich1973bredt\], Molecular Graph Validity\[polykovskiy2020molecular\]\)\. We select these filters as representative examples of commonly\-used filters in conventional drug discovery\[mignani2018present,kralj2023molecular\], removing undesirable chemical moieties and capturing a variety of perspectives on the properties of drug\-like molecules collected by chemists over decades of research\. Please note, these filters are not absolute determinants of success in drug discovery, but provide intuitive guidelines that reflect the real\-world experience of medicinal chemists\. Filter pass rates of marketed drugs are also reported, providing a fair external benchmark for general medicinal chemistry quality\. Note that these filter pass rates differ from the property metrics of Table[4](https://arxiv.org/html/2607.12349#Sx4.T4): Filter pass rates estimate rule\-based drug\-likeness and screen for known alert substructures, whereas Table[4](https://arxiv.org/html/2607.12349#Sx4.T4)verifies the chemical validity of molecules independent of drug\-likeness requirements\.
Ligands generated by๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsand๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitspass each filter at rates competitive with or higher than marketed drugs, indicating that they mirror the properties of approved drugs\. This suggests the generated ligands are largely compatible with conventional drug discovery wisdom, as reflected by these filters and rules\. Notably,๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsimproves the pass rates over๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitson pharmaceutical structural alerts \(BMS, PAINS, SureChEMBL, NIBR\), demonstrating its capability to reduce undesirable toxicophoric and reactive motifs through ADMET\-oriented optimization\. Overall, these filters highlight the promising nature of ligands generated from๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limits, while showing that๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitscan generally improve or maintain this performance, avoiding undesirable chemical moieties without harming the overall quality of generated ligands\.
### Analysis of Atom and Bond Composition on๐ข๐ฃ๐ง\\mathop\{\\mathsf\{CDH\}\}\\limits
To evaluate the capability of the models in generating realistic molecular structures, it is essential to analyze the composition of atoms and bonds in generated ligands for chemical feasibility\. As shown in Figure[4](https://arxiv.org/html/2607.12349#Sx4.F4), in general,๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsand๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsproduce atom and bond compositions that are aligned with the reference distribution in training data of๐ข๐ฃ\\mathop\{\\mathsf\{CD\}\}\\limits\. Compared with the reference, all models show a slight overproduction of carbon\. This occurs because carbon dominates the training data \(67\.2%\), so the models trained on it tend to sample it more frequently\. Consequently, other atom types are slightly under sampled\. This pattern also appears in the distribution of bond types, where single bonds occur more often than in the training data \(52\.4%\)\. Among all bond types, aromatic rings require the most strict topological and geometric constraints\. Diffusion\-based models jointly denoise atom positions and types by focusing on local rather than global structures, which makes them prone to losing fine\-grained substructures, such as aromatic rings\. In contrast,๐ฏ๐๐ผ๐๐พ๐๐ค๐ฌ๐๐
\\mathop\{\\mathsf\{Pocket2Mol\}\}\\limitspredicts atoms and bonds sequentially conditioned on previously sampled structure\. The bond prediction and the dependencies on previously generated structures allow it to better capture aromatic patterns\.
Fig\. 4:Comparison of atom \(a\-dfor C, N, O, and others \(F, S, Cl, P\), respectively\) and bond \(e\-hfor Single, Double, Trible, Aromatic, respectively\) type distributions across different models\. The dashed line displays the percentages in the training data as a realistic reference for drug molecules\.๐จ๐ฏ๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{IPDiff\}\}\\limitsachieves the best predicted binding affinity scores, as reported in TableLABEL:tbl:overall\_results\_crossdockand Table[1](https://arxiv.org/html/2607.12349#Sx4.T1)\. Its high affinity scores may stem from generating carbon\-heavy, single\-bond\-dominated structures with few heteroatoms and limited aromaticity, as Figure[4](https://arxiv.org/html/2607.12349#Sx4.F4)shows\. Excessive carbons in the ligand scaffold render the ligand broadly hydrophobic and prone to bind the pocket in a non\-specific way\. Although this may yield high Vina scores, it does not indicate a high intrinsic binding affinity\. The hydrophobic behavior also implies high lipophilicity, which often leads to poor ADMET properties and a higher risk of off\-target interactions\. At the same time, the scarcity of heteroatoms further lowers solubility and selectivity\. Therefore, although the ligands generated by๐จ๐ฏ๐ฃ๐๐ฟ๐ฟ\\mathop\{\\mathsf\{IPDiff\}\}\\limitsachieve high Vina scores, they are less attractive as drug candidates\.
### Case Studies
To demonstrate the practical utility of our methods, we use the programmed death\-ligand 1 \(PD\-L1\) and the colony\-stimulating factor\-1 receptor \(CSF1R\) targets for case studies\. PD\-L1 is a critical target for cancer immunotherapy due to its pivotal role in immune evasion by tumors\[mandal2025overcoming\]\. Current anti\-PD\-L1 therapies are being hampered by several clinical limitations, such as modest efficacy and resistance\[mandal2025overcoming\]\. Several high\-resolution crystal structures of PD\-L1\-ligand complexes are publicly available in the PDB database\[berman2000protein,Guzik2017\]\. CSF1R is a biologically important kinase involved in macrophage survival, proliferation, and differentiation\[wen2023csf1r,cannarile2017colony\]\. Current CSF1R inhibitors on the market, for example, pexidartinib and vimseltinib, have emerged as safe and efficacious options for the treatment of tenosynovial giant cell tumor \(TGCT\), a non\-malignant tumor of the joint, tendon sheath, or bursa driven by the overexpression of CSF1\. However, approved CSF1R\-targeting drugs still remain limited, motivating the development of additional candidate inhibitors\. CSF1R has a publicly available high\-quality co\-crystal structure\[kane2024identification\]\.
#### Experimental Setup
For both PD\-L1 and CSF1R, the case study pipeline consists of two sequential and complementary stages:\(1\)a structure\-based design stage to generate candidate ligands by๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits, and\(2\)a ligand\-based design stage to expand the๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated ligand set through analog search in existing corporate compounds deck and to prioritize promising compounds, thereby mimicking a hit expansion campaign as routinely done in early drug discovery\. These two stages together provide a rigorous workflow for candidate generation, screening, and selection prior to synthesis and subsequent biological testing\.
Fig\. 5:Docked poses of compound PL\-0 \(shown in magenta\), a known cognate ligand of PD\-L1, and compound PL\-1 \(shown in green\), a๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated novel ligand of PD\-L1, suggest that both compounds exhibit similar binding modes\.
Fig\. 6:Docked poses of compound CL\-0 \(shown in magenta\), a known cognate ligand of CSF1R, and compound CL\-12 \(shown in green\), a๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated novel ligand of CSF1R, suggest that both compounds exhibit similar binding modes\.
##### Structure\-Based Design via๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits
The co\-crystal structures of PD\-L1 and CSF1R bound to one of their respective cognate ligands are retrieved from the PDB\[berman2000protein\]under code 5N2F\[Guzik2017\]and 8W1L\[pdb8w1l,kane2024identification\], respectively\. The targets are prepared at pH 7\.4 using the protein wizard module as implemented in Maestro\[maestro2026\]\. Top\-ranked ligands generated by๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsin both cases were further filtered to remove undesirable chemical moieties as previously described by Kombo*et al\.*\[kombo2024predictions\]\. The remaining ligand structures are protonated at pH 7\.4 and energy\-minimized using Ligprep as implemented in Maestro\[maestro2026\]\. Flexible ligand docking was carried out using GLIDE in standard precision mode \(SP\), followed by extra precision mode \(XP\)\.
##### Ligand\-Based Design based on๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsGeneration
Following the structure\-based design stage, ligand\-based virtual screening of the Sanofi compound collection was carried out using FastROCS\[grant1996fast,rush2005shape,hawkins2007comparison\], a GPU\-accelerated tool for shape similarity search\. Ligands generated by๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsare used to build the shape query\. To further prioritize the identified virtual hits, machine learning \(ML\) models developed using structure\-activity relationship \(SAR\) datasets experimentally\-determined at Sanofi\[kane2024identification\]are used to predict potency against PD\-L1 and CSF1R, respectively\[kombo2024predictions\]\. Derived analogs are further scored and ranked using a weighted Pareto multi\-parameter optimization approach combining docking, ML\-predicted potency, and 3D similarity in shape and chemical features\.
#### Docking Results
Binding modes of reference compounds,๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated molecules, and their analogs are predicted by docking using GLIDE\[maestro2026\]\. Figure[6](https://arxiv.org/html/2607.12349#Sx4.F6)and[6](https://arxiv.org/html/2607.12349#Sx4.F6)show examples for PD\-L1 and CSF1R, respectively\. In the case of PD\-L1, the terminal phenyl group of the quintessential bi\-phenyl moiety in both the reference ligand and the ligand generated by๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsmakeฯ\\pi\-ฯ\\piinteractions with TYR\-56\. The other end of both ligands interacts with ASP\-122, one of the key residues driving molecular recognition at the PD\-L1 dimer interface\. In the case of CSF1R, both reference and๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated ligand make hydrogen\-bond interactions with CYS\-666 and ASP\-796, which constitute a hallmark of kinase inhibitors of this class\. These examples indicate that๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated ligands and their corresponding reference ligands have similar binding modes, further emphasizing the ability of๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsto learn from the binding site characteristics and accordingly generate high\-quality binding ligands\.
#### Biological Testing Results
PL\-0PL\-1PL\-2PL\-3PL\-4Fig\. 7:PD\-L1 ligand structures\. Each structure is labeled by its ID\.Table 5:SPR\-derived binding affinity and molecular properties of๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated PD\-L1 ligands and their analogs\.IDSourceKD\{\}\_\{\\text\{D\}\}\(ฮผ\\muM\)โ\\downarrowLogDQEDโ\\uparrowLipEโ\\uparrowPL\-1๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits3\.493\.010\.652\.45PL\-2๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits3\.753\.010\.652\.42PL\-3๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsanalog \(modified\)4\.002\.700\.642\.70PL\-4๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsanalog \(modified\)12\.300\.330\.504\.58PL\-0PDB: 5N2FND3\.040\.43ND
- โขโโColumns represent: โSourceโ: the method used to obtain the molecule structure โ โmodifiedโ indicates the analog was manually modified from๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated molecules; โKD\{\}\_\{\\text\{D\}\}โ: SPR\-derived equilibrium dissociation constant inฮผ\\muM; โLogDโ: distribution coefficient, with a range between 1\-3 generally considered optimal in drug discovery; โQEDโ: quantitative estimate of drug\-likeness; โLipEโ: lipophilic efficiency;โ\\uparrow/โ\\downarrowindicates that higher / lower values are better\.
CL\-0CL\-1CL\-2CL\-3CL\-4CL\-5CL\-6CL\-7CL\-8CL\-9CL\-10CL\-11CL\-12Fig\. 8:CSF1R ligand structures\. Each structure is labeled by its ID\.Table 6:Kinase activity, molecular properties, and promiscuity profiling of๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated CSF1R molecules and their analogs\.IDSourceIC50\(nM\)โ\\downarrowLogDQEDโ\\uparrowLipEโ\\uparrow\#TargetsAssayed\#Target Hits\(ICโค5010ฮผ\{\}\_\{50\}\\leq 10~\\muM\)PromiscuityHit Rate \(%\)โ\\downarrowCL\-0PDB:8W1L201\.920\.365\.7836411\.1CL\-1๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsanalog \(VS\)2005\.060\.431\.649111\.1CL\-2๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsanalog \(VS\)3095\.750\.320\.7712632\.4CL\-3๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsanalog \(VS\)2,5005\.050\.460\.5512700CL\-4๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsanalog \(VS\)3,3632\.690\.482\.7917531\.7CL\-5๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsanalog \(VS\)3,7362\.140\.693\.298900CL\-6๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsanalog \(VS\)5,9391\.870\.423\.3511200CL\-7๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsanalog \(VS\)11,7002\.350\.632\.589500CL\-8๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsanalog \(VS\)21,2570\.930\.433\.742700CL\-9๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsanalog \(VS\)\>โ30,000\\mathord\{\>\}30,0001\.240\.48ND9400CL\-10๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsInactive2\.820\.69NDโโโCL\-11๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsInactive4\.820\.75NDโโโCL\-12๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsND4\.660\.48NDโโโ
- โขโโColumns represent: โSourceโ: the method used to obtain the molecule structure, where โVSโ indicates the analog was identified via virtual screening with๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated molecules as shape query; โIC50\(nM\)โ: kinase activity against CSF1R measured in nM; โLogDโ: distribution coefficient, with a range between 1โ3 generally considered optimal in drug discovery; โQEDโ: quantitative estimate of drug\-likeness; โLipEโ: lipophilic efficiency; โ\#Targets Assayedโ: the number of distinct targets assayed in the selectivity or promiscuity panel; โ\#Target Hitsโ: the number of targets with ICโค5010ฮผ\{\}\_\{50\}\\leq 10~\\muM; โPromiscuity Hit Rate \(%\)โ: the percentage of assayed targets with ICโค5010ฮผ\{\}\_\{50\}\\leq 10~\\muM; โInactiveโ indicates no significant inhibition was observed; โNDโ indicates not determined;โ\\uparrow/โ\\downarrowindicates that higher / lower values are better\.
Four compounds generated by๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\(PL\-1, PL\-2, CL\-10, CL\-11\), nine analogs derived from ligand\-based virtual screening of the Sanofi Corporate compound collection \(CL\-1, CL\-2, CL\-3, CL\-4, CL\-5, CL\-6, CL\-7, CL\-8, CL\-9\), and two manually modified analogs \(PL\-3, PL\-4\) have been synthesized \(AppendixLABEL:supp:synthesispresents the synthesis process\) and tested in biological assays \(AppendixLABEL:supp:assaypresents the biological testing protocols\)\. Lipophilic ligand efficiency \(LipE\) is calculated by subtracting calculated LogD at pH 7\.4 from pIC50and pKD\{\}\_\{\\text\{D\}\}in the case of CSF1R and PD\-L1 ligands, respectively\. LogD and QED were calculated using Pipeline Pilot\[pilot2020version\]\. The biological activity data for PD\-L1 and CSF1R are summarized in Tables[5](https://arxiv.org/html/2607.12349#Sx4.T5)and[6](https://arxiv.org/html/2607.12349#Sx4.T6), respectively; compound structures are presented in Figures[7](https://arxiv.org/html/2607.12349#Sx4.F7)and[8](https://arxiv.org/html/2607.12349#Sx4.F8), respectively\. AppendixLABEL:supp:ic50presents the dose\-response curves for CSF1R\.
For PD\-L1, as shown in Table[5](https://arxiv.org/html/2607.12349#Sx4.T5), the๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated ligands \(PL\-1, PL\-2\) and one of their modified analogs \(PL\-3\) exhibit low\-micromolar binding affinity, with values ranging from 3\.49 to 4\.00ฮผ\\muM and LipE values ranging from 2\.42 to 2\.70\. The other modified analog, PL\-4, exhibits binding at 12\.30ฮผ\\muM with a LipE of 4\.58\. It is encouraging to note that published mean LipE value for oral drugs is 4\.43\[hopkins2014role\], and LipE values increase from hit finding to drug candidate discovery\[kombo2025logic\]โ Some๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated ligands already achieve ligand efficiency profiles comparable to established orally available drugs\. All four novel ligands exhibit QED values greater than the reference ligand \(PL\-0\)\.
As shown in Table[6](https://arxiv.org/html/2607.12349#Sx4.T6), results indicate that using๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated ligands as a template for building a 3D search query to find CSF1R ligands yielded two compounds, CL\-1 and CL\-2, with nanomolar kinase activity against CSF1R, with IC50values of 200 nM and 309 nM, respectively, and four other compounds, CL\-3, CL\-4, CL\-5, and CL\-6, with low micromolar kinase activity, with IC50values ranging from 2\.500 to 5\.939ฮผ\\muM\. Moreover, these compounds show low kinase promiscuity across the tested panels\. Specifically, under the 10ฮผ\\muM threshold, CL\-1 and CL\-2 have promiscuity hit rates of only 1\.1% and 2\.4%, respectively, which are substantially lower than the 11\.1% observed for the reference ligand \(CL\-0\)\. Among the other four compounds, all except CL\-4 show no detectable off\-target kinase hits under the same threshold, while CL\-4 shows only a low promiscuity hit rate of 1\.7%\. In addition, some of the virtual screening hits are known Factor Xa inhibitors, as shown by an expired patent\[ewing2001substituted\], which can inspire leveraging the observed polypharmacology to initiate compound repositioning studies\. Taken together, the results obtained in both target cases suggest that๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated ligands and their subsequent analogs have developability characteristics like those found in hit finding and hit expansion campaigns as carried out in drug discovery projects in pharmaceutical settings\. Furthermore, we have predicted that๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-generated ligands and their corresponding reference ligands have similar binding modes, further emphasizing the ability of๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitswhen it comes to efficiently deciphering the binding site characteristics\.
## Discussions and Conclusions
In this section, we discuss the limitations of๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsand๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsand sketch promising directions for future work\.
### Specializing Model for High\-selectivity Ligands
An important consideration in SBDD is ligand selectivity\. While current diffusion\-based models focus primarily on optimizing binding affinity to a given target, they often ignore off\-target interactions, which can lead to undesirable side effects\. A ligand that binds to both the intended target and unintended proteins may exhibit poor selectivity, reducing its therapeutic potential and increasing the risk of adverse effects\. To address this, future work could explore strategies for specializing generative models to generate highly selective ligands\. One approach is negative conditioning, training the model to avoid off\-targets while retaining on\-target affinity\. Alternatively, optimization techniques over on\- and off\-target scores can balance affinity and selectivity during generation\. These computational strategies could potentially improve the clinical relevance of structure\-based generated ligands and advance the applicability of such AI\-generative models in drug discovery\.
### Generalizing Model to Multi\-objective Optimization
While our current generation\-time optimization targets a single ADMET property at a time, real discovery tasks generally require simultaneous consideration of multiple ADMET endpoints, since strong affinity alone often fails to translate when even one critical ADMET criterion is not met\. A natural extension is a multi\-objective formulation that integrates multiple property signals simultaneously into the optimization module\. Such a method may construct composite update directions that coordinate improvements across all target properties or dynamically adapt the relative emphasis of each property according to which constraints are currently violated\. This would move the current single\-objective setting toward Pareto\-style multi\-objective optimization and enable fine\-grained control over affinity and ADMET trade\-offs while preserving the training\-free and plug\-and\-play nature of the optimization module\. For instance, AI\-MedCraft demonstrates how an adaptive Pareto\-guided strategy can coordinate multiple competing objectives\[barakat2026ai\]\. In practice, such flexibility could support different stages of discovery\. Tighter ADMET constraints could reduce downstream risk when optimizing near\-deployable candidates, whereas looser constraints could encourage broader exploration in early ideation, followed by gradual annealing toward stricter feasibility as optimization progresses\. Another promising direction is to make the optimization uncertainty\-aware\. Property predictors are often noisy, susceptible to distribution shift, and potentially vulnerable to exploitation under direct optimization\. To address this, the objective could incorporate risk\-sensitive criteria, such as lower\-confidence\-bound optimization or explicit penalties on predictive uncertainty, thereby steering updates toward candidates that are not only high\-scoring but also more reliable\. Together, these extensions would better align generative optimization with pharmacological and drug\-discovery practice, potentially accelerating ligand optimization toward clinic\-viable candidates\.
### Incorporating Chemical Synthesis Pathway into Diffusion Models
A major practical concern in drug discovery is the synthetic accessibility of proposed ligands\. While conDitar and conDitar\-dev generate chemically valid molecules and optimize the ADMET properties, they do not explicitly account for whether the binding ligand can be realized through feasible synthetic pathways\. As a result, some generated ligands, despite exhibiting good*in silico*profile, may be difficult to synthesize\. To address this limitation, future work could explore incorporating synthesis pathway information into diffusion models\. One possible direction is to introduce synthesis\-aware priors or guidance signals, such as retrosynthetic accessibility scores or reaction constraints, to bias generation toward synthetically accessible regions of chemical space\[shen2025compositional,cremer2024pilot\]\. Such a space could include commercially available chemical building blocks commonly used in chemical reactions carried out by medicinal chemists\. Alternatively, synthetic pathways could be modeled explicitly during diffusion training, enabling the diffusion process to simultaneously simulate molecular generation and the synthesis trajectory\. Such synthesis\-aware generative frameworks could enhance the alignment between generative model outputs and practical requirements in real\-world drug discovery workflows\.
### Informing Diffusion Models with Physical and Chemical Knowledge
Diffusion\-based methods may occasionally generate molecules with atypical ring systems, achieving high binding affinity scores while raising concerns regarding chemical realism\. To address this, one possible direction is to incorporate chemical domain knowledge and physical constraints into the generative process, thereby constructing a chemistry\- and physics\-informed diffusion framework\. For example, a framework like๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitscould be coupled with force field\-based energy minimization or other physics\-based refinement procedures to improve the structural realism of the generated ligands\. One such force field\-based method is OPLS\-AA\[jorgensen1988opls,jorgensen1996development,sambasivarao2009development,doherty2017revisiting\], which is widely used in molecular modeling and simulations of interactions between organic molecules and proteins\. Furthermore, in future work, the generated ligands and pocket residue side chains should be properly protonated at physiological pH \(7\.4\) prior to energy minimization and docking to ensure physically meaningful interaction modeling\. In addition to these improvements, we envision future work incorporating a robust protocol aimed at concomitantly filtering undesirable chemical moieties as ligands are being generated to steer the design process towards a more druggable chemical space\.
### Conclusions
We present๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits, a conditional diffusion\-based framework for SBDD capable of generating ligands with strong binding affinities and ADMET properties for drug discovery and development\. On๐ข๐ฃ๐ง\\mathop\{\\mathsf\{CDH\}\}\\limits,๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsachieves better performance than the state\-of\-the\-art SBDD baselines, achieving an average Vina D score of \-8\.85 kcal/mol\. On five ADMET properties,๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsachieves improvements of up to 73% over๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitswhile retaining similar predicted binding affinity\. Furthermore, the molecules generated by๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsand๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsexhibit strong stability scores, realistic chemical structures, and high drug discovery filter pass rates\. Together, these results establish๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsas a strong framework for generating developable drug molecules\. Moreover,๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\-designed molecules for CSF1R and PD\-L1, together with their derivatives and analogs, have been experimentally synthesized and biologically tested\. Several compounds show promising binding affinities in the nanomolar and low\-micromolar range, validating๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsbeyond the typical standard for*in silico*SBDD frameworks, which typically lack experimental validation\[hu2025target\]\. This case study demonstrates how๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitscan be used immediately by computational and medicinal chemists, highlighting its potential to transform drug discovery\. We envision๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsas a foundation for more advanced SBDD frameworks that incorporate additional developability objectives, including expanded ADMET properties, synthetic accessibility, structural constraints, and ligand selectivity\.
## Methods
We introduce๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits, a new framework for generating 3D ligands tailored to specific protein pockets, while also optimizing additional ADMET properties of the generated ligands\. As illustrated in Figure[1](https://arxiv.org/html/2607.12349#Sx1.F1),๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitscomprises three modules:\(1\)๐๐๐ฏ๐ฑ๐ซ\\mathop\{\\mathsf\{msPRL\}\}\\limits, a pretrained pocket embedding module,\(2\)๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limits, a diffusion\-based, pocket\-conditioned ligand generation module, and\(3\)๐๐บ๐ฎ๐ฏ๐ณ\\mathop\{\\mathsf\{paOPT\}\}\\limits, an optional optimization module applied at generation time of๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limits\.๐๐๐ฏ๐ฑ๐ซ\\mathop\{\\mathsf\{msPRL\}\}\\limitslearns to represent 3D structures and physicochemical properties of protein binding pockets through an equivariant encoder\. For a given pocket, referred to as the condition pocket, its binding pocket representation produced by๐๐๐ฏ๐ฑ๐ซ\\mathop\{\\mathsf\{msPRL\}\}\\limitswill be used in diffusion to generate new ligands\.๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsgenerates new, realistic ligands in 3D with high binding affinities to the condition pocket by learning from the pocket\-ligand interactions from existing complexes via a diffusion process\[ho2020ddpm\]\. To further improve other desired pharmacological properties of the generated ligands \(e\.g\., ADMET properties\[hodgson2001admet\]\),๐๐บ๐ฎ๐ฏ๐ณ\\mathop\{\\mathsf\{paOPT\}\}\\limitsis built upon๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsby performing gradient\-based, zero*th*\-order optimization during its generation time without retraining๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limits, guiding the diffusion trajectories toward the desired properties\. Table[7](https://arxiv.org/html/2607.12349#Sx6.T7)presents the key notations used in the methods\. Particularly, subscriptsiandjindex atoms or residues\. Superscripts indicate the data domain:gfor ligand atoms,pfor pocket atoms,rfor pocket residues,cfor interactions between ligand atoms and pocket atoms, andafor interactions between ligand atoms and pocket residues\.
Table 7:Notationsnotationsmeanings๐\\mathop\{\\mathcal\{D\}\}\\limitsa ligand๐ซ\\mathop\{\\mathcal\{P\}\}\\limitsa pocket๐บi\\mathsf\{a\}\_\{i\}theii\-th atom in๐\\mathop\{\\mathcal\{D\}\}\\limits/๐ซ\\mathop\{\\mathcal\{P\}\}\\limits๐ip\\mathsf\{r\}^\{p\}\_\{i\}theii\-th residue in๐ซ\\mathop\{\\mathcal\{P\}\}\\limits๐ฑi\\mathbf\{x\}\_\{i\}the position of๐บi\\mathsf\{a\}\_\{i\}/๐ip\\mathsf\{r\}^\{p\}\_\{i\}in 3D space๐ฏi\\mathbf\{v\}\_\{i\}the atom/residue type of๐บi\\mathsf\{a\}\_\{i\}/๐ip\\mathsf\{r\}^\{p\}\_\{i\}๐ฌi\\mathbf\{s\}\_\{i\}scalar embedding for atom๐บi\\mathsf\{a\}\_\{i\}โi\\mathcal\{H\}\_\{i\}vector embedding for atom๐บi\\mathsf\{a\}\_\{i\}dโ\(๐บi,๐บj\)d\(\\mathsf\{a\}\_\{i\},\\mathsf\{a\}\_\{j\}\)the distance between๐บi\\mathsf\{a\}\_\{i\}and๐บj\\mathsf\{a\}\_\{j\}๐iโj\\mbox\{$\\mathop\{\\mathbf\{b\}\}\\limits$\}\_\{ij\}the type of bond between๐บig\\mathsf\{a\}^\{g\}\_\{i\}and๐บjg\\mathsf\{a\}^\{g\}\_\{j\}gligand atomsppocket atomsrpocket residuescinteractions between ligand atoms and pocket atomsainteractions between ligand atoms and pocket residues### Preliminaries
Given the condition protein pocket๐ซ\\mathop\{\\mathcal\{P\}\}\\limits,๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsโs objective is to generate ligands exhibiting strong binding affinities to๐ซ\\mathop\{\\mathcal\{P\}\}\\limitswith realistic 3D structures\.๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsextends this objective by optimizing additional physicochemical properties\. In this manuscript, we represent a ligand๐\\mathop\{\\mathcal\{D\}\}\\limitsas a set of its atoms, defined as
๐=\{๐บ1g,๐บ2g,โฆ,๐บNgโฃ๐บig=\(๐ฑig,๐ฏig\)\},\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}=\\\{\\mathsf\{a\}^\{g\}\_\{1\},\\mathsf\{a\}^\{g\}\_\{2\},\\dots,\\mathsf\{a\}^\{g\}\_\{N\}\\mid\\mathsf\{a\}^\{g\}\_\{i\}=\(\\mathbf\{x\}^\{g\}\_\{i\},\\mathbf\{v\}^\{g\}\_\{i\}\)\\\},whereNNdenotes the number of atoms in๐\\mathop\{\\mathcal\{D\}\}\\limits\. Each atom๐บig\\mathsf\{a\}^\{g\}\_\{i\}is characterized by its 3D coordinates๐ฑigโโ1ร3\\mathbf\{x\}^\{g\}\_\{i\}\\in\\mathbb\{R\}^\{1\\times 3\}and a one\-hot feature vector๐ฏigโโ1รda\\mathbf\{v\}^\{g\}\_\{i\}\\in\\mathbb\{R\}^\{1\\times\{d\_\{a\}\}\}\. This feature vector encodes both the atom type and its aromaticity\. Here,da=13d\_\{a\}=13denotes the total number of distinct atom types considered \(e\.g\., carbon, oxygen, nitrogen; hydrogen is excluded\), including aromatic and non\-aromatic variants\. Similarly, we represent a protein pocket๐ซ\\mathop\{\\mathcal\{P\}\}\\limitsusing its set of protein atoms within a predefined radius of its reference ligand \(as introduced in the section Datasets\) denoted asAp\\mathop\{\{A\}\_\{p\}\}\\limits, and the residues involved in the protein pocket, denoted asRp\\mathop\{\{R\}\_\{p\}\}\\limits, that is,
๐ซ=\(Ap,Rp\),\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}=\(\\mbox\{$\\mathop\{\{A\}\_\{p\}\}\\limits$\},\\mbox\{$\\mathop\{\{R\}\_\{p\}\}\\limits$\}\),where
Ap=\{๐บ1p,๐บ2p,โฏ,๐บNap\|๐บip=\(๐ฑip,๐ฏip\)\},\\displaystyle=\\\{\\mathsf\{a\}^\{p\}\_\{1\},\\mathsf\{a\}^\{p\}\_\{2\},\\cdots,\\mathsf\{a\}^\{p\}\_\{N\_\{a\}\}\|\\mathsf\{a\}^\{p\}\_\{i\}=\(\\mathbf\{x\}^\{p\}\_\{i\},\\mathbf\{v\}^\{p\}\_\{i\}\)\\\},\(1\)Rp=\{๐1p,๐2p,โฏ,๐Nrp\|๐ip=\(๐ฑir,๐ฏir\)\}\.\\displaystyle=\\\{\\mathsf\{r\}^\{p\}\_\{1\},\\mathsf\{r\}^\{p\}\_\{2\},\\cdots,\\mathsf\{r\}^\{p\}\_\{N\_\{r\}\}\|\\mathsf\{r\}^\{p\}\_\{i\}=\(\\mathbf\{x\}^\{r\}\_\{i\},\\mathbf\{v\}^\{r\}\_\{i\}\)\\\}\.Each pocket atom๐บipโAp\\mathsf\{a\}^\{p\}\_\{i\}\\in A\_\{p\}is characterized by its 3D coordinates๐ฑipโโ1ร3\\mathbf\{x\}^\{p\}\_\{i\}\\in\\mathbb\{R\}^\{1\\times 3\}and a one\-hot feature vector๐ฏipโโ1รda\\mathbf\{v\}^\{p\}\_\{i\}\\in\\mathbb\{R\}^\{1\\times\{d\_\{a\}\}\}, representing its atom type withda=4d\_\{a\}=4\(carbon, nitrogen, oxygen, or sulfur\)\. Similarly, each pocket residue๐ipโRp\\mathsf\{r\}^\{p\}\_\{i\}\\in R\_\{p\}is represented using the 3D coordinates of its alpha carbon atom๐ฑirโโ1ร3\\mathbf\{x\}^\{r\}\_\{i\}\\in\\mathbb\{R\}^\{1\\times 3\}, and a one\-hot feature vector๐ฏirโโ1ร20\\mathbf\{v\}^\{r\}\_\{i\}\\in\\mathbb\{R\}^\{1\\times\{20\}\}, representing the 20 amino acid types\. Note that each residue can be decomposed into its component atoms, represented as
๐ip=\{๐บi,1p,โฆ,๐บi,\|๐ip\|p\},\\mathsf\{r\}^\{p\}\_\{i\}=\\\{\\mathsf\{a\}^\{p\}\_\{i,1\},\\dots,\\mathsf\{a\}^\{p\}\_\{i,\|\\mathsf\{r\}^\{p\}\_\{i\}\|\}\\\},where๐บi,jp\\mathsf\{a\}^\{p\}\_\{i,j\}is thejj\-th atom in๐ip\\mathsf\{r\}^\{p\}\_\{i\}\(๐บi,jpโAp\\mathsf\{a\}^\{p\}\_\{i,j\}\\in A\_\{p\}\), and\|๐ip\|\|\\mathsf\{r\}^\{p\}\_\{i\}\|is the number of atoms in๐ip\\mathsf\{r\}^\{p\}\_\{i\}\. For each๐บjpโAp\\mathsf\{a\}^\{p\}\_\{j\}\\in A\_\{p\}, the residue that it is in is denoted as๐pโ\(๐บjp\)\\mathsf\{r\}^\{p\}\(\\mathsf\{a\}^\{p\}\_\{j\}\), that is,๐pโ\(๐บi,jp\)=๐ip\\mathsf\{r\}^\{p\}\(\\mathsf\{a\}^\{p\}\_\{i,j\}\)=\\mathsf\{r\}^\{p\}\_\{i\}\.
### Pocket Representation Pretraining \(๐๐๐ฏ๐ฑ๐ซ\\mathop\{\\mathsf\{msPRL\}\}\\limits\)
We pretrain a pocket embedding module, denoted as๐๐๐ฏ๐ฑ๐ซ\\mathop\{\\mathsf\{msPRL\}\}\\limits, to encode the 3D structures of target protein pockets into pocket embeddings\.๐๐๐ฏ๐ฑ๐ซ\\mathop\{\\mathsf\{msPRL\}\}\\limitsemploys an encoder,๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits, which maps each pocket atom๐บip\\mathsf\{a\}^\{p\}\_\{i\}into two latent representations: a scalar embedding๐ฌipโโd\\mathbf\{s\}^\{p\}\_\{i\}\\in\\mathbb\{R\}^\{d\}and a vector embeddingโipโโdร3\\mathcal\{H\}^\{p\}\_\{i\}\\in\\mathbb\{R\}^\{d\\times 3\}, whereddis a hyperparameter of๐๐๐ฏ๐ฑ๐ซ\\mathop\{\\mathsf\{msPRL\}\}\\limits\. Together, these latent representations capture the essential structural information of the pocket\. Simultaneously,๐๐๐ฏ๐ฑ๐ซ\\mathop\{\\mathsf\{msPRL\}\}\\limitsemploys a decoder,๐๐๐ฏ๐ฑ๐ซ\-โ๐ฝ๐พ๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{dec\}\}\\limits, to decode pocket embeddings to atom positions and types\. FollowingUni\-Molโs\[zhou2023uni\]approach,๐๐๐ฏ๐ฑ๐ซ\\mathop\{\\mathsf\{msPRL\}\}\\limitsis optimized by reconstructing noise\-free pockets from corrupted ones using pocket embeddings, where the pockets are corrupted by adding uniform noise into the positions of randomly selected atoms with their atom types masked off\. Existing diffusion\-based methods\[guan2023targetdiff,guan2023decompdiff,huang2024protein\]typically jointly learn pocket representations and ligand representations together from pocket\-ligand interactions\. This approach can limit the quality of pocket representation learning since the comparatively larger size of pockets makes them more prone to being underrepresented or traded off relative to ligands\. In๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limits, we isolate the pocket representation learning within the pretraining of the๐๐๐ฏ๐ฑ๐ซ\\mathop\{\\mathsf\{msPRL\}\}\\limitsmodule to ensure expressive pocket embeddings\.
#### Pocket Encoder \(๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits\)
๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limitslearns a scalar embedding๐ฌip\\mathbf\{s\}^\{p\}\_\{i\}and a vector embeddingโip\\mathcal\{H\}^\{p\}\_\{i\}for each atom๐บip\\mathsf\{a\}^\{p\}\_\{i\}in the pocket๐ซ\\mathop\{\\mathcal\{P\}\}\\limitsto encode pocket structure and physicochemical properties\. These two embeddings integrate the information of each atom๐บipโAp\\mathsf\{a\}^\{p\}\_\{i\}\\in A\_\{p\}and the residue๐pโ\(๐บip\)โRp\\mathsf\{r\}^\{p\}\(\\mathsf\{a\}^\{p\}\_\{i\}\)\\in\\mbox\{$\\mathop\{\{R\}\_\{p\}\}\\limits$\}to which๐บip\\mathsf\{a\}^\{p\}\_\{i\}belongs to capture multi\-granular information from the pockets\.
##### Pocket Atom Embeddings
๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limitsuses a multi\-layer graph attention neural network\[velickovic2018graph\]augmented with a geometric vector perceptron \(GVP\)\[jing2021learning\]to learn embeddings of pocket atoms\. At each layerllof the graph neural network,๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limitslearns the invariant features/embeddings๐ฌi,lp\\mathbf\{s\}^\{p\}\_\{i,l\}and the equivariant features/embeddingsโi,lp\\mathcal\{H\}^\{p\}\_\{i,l\}\(๐ฌi,0p=๐ฏip\\mathbf\{s\}^\{p\}\_\{i,0\}=\\mathbf\{v\}^\{p\}\_\{i\},โi,0p=๐ฑip\\mathcal\{H\}^\{p\}\_\{i,0\}=\\mathbf\{x\}^\{p\}\_\{i\}\)\. Invariant features capture scalar properties such as atom types and interatomic distances, which remain unchanged under rotations and translations; equivariant features encode vector information such as atom positions and spatial directions that are difference vectors between atom positions, which transform consistently under these transformations\. AfterLLlayers,๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limitsoutputs๐ฌi,Lp\\mathbf\{s\}^\{p\}\_\{i,L\}andโi,Lp\\mathcal\{H\}^\{p\}\_\{i,L\}as the final pocket atom embeddings, that is,
๐ฌip=๐ฌi,Lp,โip=โi,Lp\.\\mathbf\{s\}^\{p\}\_\{i\}=\\mathbf\{s\}^\{p\}\_\{i,L\},\\mathcal\{H\}^\{p\}\_\{i\}=\\mathcal\{H\}^\{p\}\_\{i,L\}\.\(2\)
Specifically, at thell\-th layer,๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limitsupdates the embeddings๐ฌi,lp\\mathbf\{s\}^\{p\}\_\{i,l\}andโi,lp\\mathcal\{H\}^\{p\}\_\{i,l\}for atom๐บip\\mathsf\{a\}^\{p\}\_\{i\}by aggregating information from its neighboring atoms as follows:
\(๐ฌi,lp,โi,lp\)\\displaystyle\(\\mathbf\{s\}^\{p\}\_\{i,l\},\\mathcal\{H\}^\{p\}\_\{i,l\}\)=GVPโ\(๐กi,lp,๐ดi,lp\),where\\displaystyle=\\text\{GVP\}\(\\mathbf\{h\}^\{p\}\_\{i,l\},\\mathcal\{Y\}^\{p\}\_\{i,l\}\),\\text\{ where \}\(3\)๐กi,lp\\displaystyle\\mathbf\{h\}^\{p\}\_\{i,l\}=\[๐ฏip,๐ฌi,lโ1p,โ๐บjpโ๐\(๐บip\|๐ซ\)ejโi,l๐ฆjโi,lp\],\\displaystyle=\\biggr\[\\mathbf\{v\}^\{p\}\_\{i\},\\mathbf\{s\}^\{p\}\_\{i,l\-1\},\\sum\_\{\\mathsf\{a\}^\{p\}\_\{j\}\\in\{\\mbox\{$\\mathop\{\\mathbf\{N\}\}\\limits$\}\(\\mathsf\{a\}^\{p\}\_\{i\}\|\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\)\}\}e\_\{ji,l\}\\mathbf\{m\}^\{p\}\_\{ji,l\}\\biggr\],๐ดi,lp\\displaystyle\\mathcal\{Y\}^\{p\}\_\{i,l\}=\[๐ฑip,โi,lโ1p,โ๐บjpโ๐\(๐บip\|๐ซ\)ejโi,lโMjโi,lp\],\\displaystyle=\\biggl\[\\mathbf\{x\}^\{p\}\_\{i\},\\mathcal\{H\}^\{p\}\_\{i,l\-1\},\\sum\_\{\\mathsf\{a\}^\{p\}\_\{j\}\\in\{\\mbox\{$\\mathop\{\\mathbf\{N\}\}\\limits$\}\(\\mathsf\{a\}^\{p\}\_\{i\}\|\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\)\}\}e\_\{ji,l\}M^\{p\}\_\{ji,l\}\\biggr\],whereGVPโ\(โ
\)\\text\{GVP\}\(\\cdot\)models interactions between๐กi,lp\\mathbf\{h\}^\{p\}\_\{i,l\}and๐ดi,lp\\mathcal\{Y\}^\{p\}\_\{i,l\}\. Intuitively,๐กi,lp\\mathbf\{h\}^\{p\}\_\{i,l\}aggregates the chemical characteristics of๐บip\\mathsf\{a\}^\{p\}\_\{i\}, and๐ดi,lp\\mathcal\{Y\}^\{p\}\_\{i,l\}aggregates its geometric features\.๐\(๐บip\|๐ซ\)\\mbox\{$\\mathop\{\\mathbf\{N\}\}\\limits$\}\(\\mathsf\{a\}^\{p\}\_\{i\}\|\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\)denotes the set of thekk\-nearest pocket atoms to๐บip\\mathsf\{a\}^\{p\}\_\{i\}, selected based on proximity among atoms in๐ซ\\mathop\{\\mathcal\{P\}\}\\limits;๐ฆjโi,lp\\mathbf\{m\}^\{p\}\_\{ji,l\}andMjโi,lpM^\{p\}\_\{ji,l\}correspond to the scalar and vector message embeddings, respectively, which propagate information from atom๐บjpโ๐\(๐บip\|๐ซ\)\\mathsf\{a\}^\{p\}\_\{j\}\\in\\mbox\{$\\mathop\{\\mathbf\{N\}\}\\limits$\}\(\\mathsf\{a\}^\{p\}\_\{i\}\|\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\)to๐บip\\mathsf\{a\}^\{p\}\_\{i\}; andejโi,le\_\{ji,l\}denotes the attention weight between neighbor atoms๐บjp\\mathsf\{a\}^\{p\}\_\{j\}and๐บip\\mathsf\{a\}^\{p\}\_\{i\}\. The message embeddings๐ฆjโi,lp\\mathbf\{m\}^\{p\}\_\{ji,l\}andMjโi,lpM^\{p\}\_\{ji,l\}are calculated as follows:
\(๐ฆjโi,lp,Mjโi,lp\)\\displaystyle\(\\mathbf\{m\}^\{p\}\_\{ji,l\},M^\{p\}\_\{ji,l\}\)=GVPโ\(๐ฆ^jโi,lp,M^jโi,lp\),where\\displaystyle=\\text\{GVP\}\(\\hat\{\\mathbf\{m\}\}^\{p\}\_\{ji,l\},\\hat\{M\}^\{p\}\_\{ji,l\}\),\\text\{ where \}\(4\)๐ฆ^jโi,lp\\displaystyle\\hat\{\\mathbf\{m\}\}^\{p\}\_\{ji,l\}=\[๐ฌj,lโ1p,dโ\(๐บjp,๐บip\)\],\\displaystyle=\[\\mathbf\{s\}^\{p\}\_\{j,l\-1\},\{d\(\\mathsf\{a\}^\{p\}\_\{j\},\\mathsf\{a\}^\{p\}\_\{i\}\)\}\],M^jโi,lp\\displaystyle\\hat\{M\}^\{p\}\_\{ji,l\}=\[โj,lโ1p,๐ฑjpโ๐ฑip\],\\displaystyle=\[\\mathcal\{H\}^\{p\}\_\{j,l\-1\},\{\\mathbf\{x\}^\{p\}\_\{j\}\-\\mathbf\{x\}^\{p\}\_\{i\}\]\},wheredโ\(๐บjp,๐บip\)d\(\\mathsf\{a\}^\{p\}\_\{j\},\\mathsf\{a\}^\{p\}\_\{i\}\)represents the Euclidean distance between the positions of pocket atoms๐บjp\\mathsf\{a\}^\{p\}\_\{j\}and๐บip\\mathsf\{a\}^\{p\}\_\{i\}\. The attention weightsejโi,le\_\{ji,l\}are calculated as follows:
ejโi,l\\displaystyle e\_\{ji,l\}=expโก\(Qi,lโKjโi,l\)โ๐บkpโ๐\(๐บip\|๐ซ\)expโก\(Qi,lโKkโi,l\),where\\displaystyle=\\frac\{\\exp\(Q\_\{i,l\}K\_\{ji,l\}\)\}\{\\sum\_\{\\mathsf\{a\}^\{p\}\_\{k\}\\in\{\\mbox\{$\\mathop\{\\mathbf\{N\}\}\\limits$\}\(\\mathsf\{a\}^\{p\}\_\{i\}\|\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\)\}\}\\exp\(Q\_\{i,l\}K\_\{ki,l\}\)\},\\text\{ where \}\(5\)Qi,l\\displaystyle Q\_\{i,l\}=MLPโ\(\[๐ฌi,lโ1p,โโi,lโ1pโ2\]\),\\displaystyle=\\text\{MLP\}\(\{\[\\mathbf\{s\}^\{p\}\_\{i,l\-1\},\\\|\\mathcal\{H\}^\{p\}\_\{i,l\-1\}\\\|\_\{2\}\]\}\),Kjโi,l\\displaystyle K\_\{ji,l\}=MLPโ\(\[๐ฆjโi,lp,โMjโi,lpโ2\]\),\\displaystyle=\\text\{MLP\}\(\[\\mathbf\{m\}^\{p\}\_\{ji,l\},\\\|M^\{p\}\_\{ji,l\}\\\|\_\{2\}\]\),whereโฅโ
โฅ2\\\|\\cdot\\\|\_\{2\}denotes the Euclidean \(โ2\\ell\_\{2\}\) norm of a 2D vector\. Intuitively, the attention weightejโi,le\_\{ji,l\}measures how much neighbor atom๐บjp\\mathsf\{a\}^\{p\}\_\{j\}should influence the update of๐บip\\mathsf\{a\}^\{p\}\_\{i\}based on both chemical and spatial information encoded inQi,jQ\_\{i,j\}andKjโi,lK\_\{ji,l\}, allowing the model to pay more attention to more informative neighbors\. Here,Qi,lQ\_\{i,l\}represents the query of atom๐บip\\mathsf\{a\}^\{p\}\_\{i\}, summarizing the features of atom๐บip\\mathsf\{a\}^\{p\}\_\{i\}and geometry surrounding it, whileKjโi,lK\_\{ji,l\}represents the key of neighbor๐บjp\\mathsf\{a\}^\{p\}\_\{j\}, capturing its message\-level features that combine chemical and spatial information\.
##### Pocket Residue Embeddings
๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limitsalso learns residue embeddings, denoted as๐ฌir\\mathbf\{s\}^\{r\}\_\{i\}andโir\\mathcal\{H\}^\{r\}\_\{i\}, respectively, using a graph neural network \(GNN\) similar to the one described above \(Equations[3](https://arxiv.org/html/2607.12349#Sx6.E3)to[5](https://arxiv.org/html/2607.12349#Sx6.E5)\) over residues\. Incorporating this residue information allows๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limitsto capture important physicochemical properties at the residue level, such as polarity, charge, and hydrophobicity, that are critical for determining protein\-ligand interactions\[gilson2007calculation\]\.
##### Pocket Embeddings
๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limitsintegrates pocket atom embeddings and pocket residue embeddings into a single, enriched pocket embedding, comprehensively capturing the multi\-granular spatial and physicochemical features of the protein pockets\. Particularly,๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limitscombines each pocket atom embedding and the embedding of its belonging residue through a gating mechanism, which balances the two embeddings by adaptively weighting their contributions\. Thus, the atom scalar embedding for๐บip\\mathsf\{a\}^\{p\}\_\{i\}is updated as follows:
๐ฌip\\displaystyle\\mathbf\{s\}^\{p\}\_\{i\}=gisโ๐ฌip\+\(1โgis\)โ๐ฌjr,๐jp=๐pโ\(๐บip\),where\\displaystyle=g^\{s\}\_\{i\}\\mathbf\{s\}^\{p\}\_\{i\}\+\(1\-g^\{s\}\_\{i\}\)\\mathbf\{s\}^\{r\}\_\{j\},~~\\mathsf\{r\}^\{p\}\_\{j\}=\\mathsf\{r\}^\{p\}\(\\mathsf\{a\}^\{p\}\_\{i\}\),\\text\{ where \}\(6\)gis\\displaystyle g^\{s\}\_\{i\}=ฯ\(MLP\(\[๐ฌip,๐ฌjr\]\),\\displaystyle=\\sigma\(\\text\{MLP\}\(\[\\mathbf\{s\}^\{p\}\_\{i\},\\mathbf\{s\}^\{r\}\_\{j\}\]\),wheregisg^\{s\}\_\{i\}denotes the gating weight, learned from the scalar embeddings of the pocket atom๐บig\\mathsf\{a\}^\{g\}\_\{i\}and its corresponding residue๐pโ\(๐บip\)\\mathsf\{r\}^\{p\}\(\\mathsf\{a\}^\{p\}\_\{i\}\)\. Similarly, the atom vector embedding for๐บip\\mathsf\{a\}^\{p\}\_\{i\}is updated as follows:
โip\\displaystyle\\mathcal\{H\}^\{p\}\_\{i\}=gihโโip\+\(1โgih\)โโjr,๐jp=๐pโ\(๐บip\),where\\displaystyle=g^\{h\}\_\{i\}\\mathcal\{H\}^\{p\}\_\{i\}\+\(1\-g^\{h\}\_\{i\}\)\\mathcal\{H\}^\{r\}\_\{j\},~~\\mathsf\{r\}^\{p\}\_\{j\}=\\mathsf\{r\}^\{p\}\(\\mathsf\{a\}^\{p\}\_\{i\}\),\\text\{ where \}\(7\)gih\\displaystyle g^\{h\}\_\{i\}=ฯโ\(MLPโ\(\[โโipโ2,โโjrโ2\]\)\)\.\\displaystyle=\{\\sigma\(\\text\{MLP\}\(\[\\\|\\mathcal\{H\}^\{p\}\_\{i\}\\\|\_\{2\},\\\|\\mathcal\{H\}^\{r\}\_\{j\}\\\|\_\{2\}\]\)\)\}\.This gating mechanism enables the model to dynamically integrate atom\-level and residue\-level information, leading to expressive pocket representations over pocket atoms\.
#### Pocket Decoder \(๐๐๐ฏ๐ฑ๐ซ\-โ๐ฝ๐พ๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{dec\}\}\\limits\)
To train๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits, we introduce a pocket decoder, denoted as๐๐๐ฏ๐ฑ๐ซ\-โ๐ฝ๐พ๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{dec\}\}\\limits, to facilitate self\-supervised training on a contrived noise prediction task, inspired byUni\-Mol\[zhou2023uni\]\. To create supervision signals, we corrupt each pocket by applying two types of perturbations:\(1\)atom position perturbation โ uniform noises are added to the positions of a randomly selected subset of atoms, and\(2\)atom type masking โ the same subset of pocket atoms as in the position perturbing step has their atom types masked off\. The corrupted pocket is encoded to embeddings by๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits, as described in section Pocket Encoder \(๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits\)\.
๐๐๐ฏ๐ฑ๐ซ\-โ๐ฝ๐พ๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{dec\}\}\\limitslearns to recover the original atom positions and types from the embeddings of noisy pockets provided by๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits\. Specifically,๐๐๐ฏ๐ฑ๐ซ\-โ๐ฝ๐พ๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{dec\}\}\\limitspredicts the normal noise๐ง^ip\\hat\{\\mathbf\{n\}\}^\{p\}\_\{i\}added to the position of a pocket atom๐บip\\mathsf\{a\}^\{p\}\_\{i\}using the following formulation:
๐ง^ip\\displaystyle\\hat\{\\mathbf\{n\}\}^\{p\}\_\{i\}=โ๐บjpโ๐\(๐บip\|๐ซ\)cjโiโ\(๐ฑjpโ๐ฑip\)Ni,where\\displaystyle=\\sum\_\{\\mathsf\{a\}^\{p\}\_\{j\}\\in\{\\mbox\{$\\mathop\{\\mathbf\{N\}\}\\limits$\}\(\\mathsf\{a\}^\{p\}\_\{i\}\|\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\)\}\}\\frac\{c\_\{ji\}\(\{\\mathbf\{x\}^\{p\}\_\{j\}\-\\mathbf\{x\}^\{p\}\_\{i\}\}\)\}\{N\_\{i\}\},\\text\{ where \}\(8\)cjโi\\displaystyle\\quad c\_\{ji\}=MLPโ\(\[๐ฌip,๐ฌjp,dโ\(๐บjp,๐บip\),โโjpโโipโ2\]\),\\displaystyle=\\text\{MLP\}\(\[\\mathbf\{s\}^\{p\}\_\{i\},\\mathbf\{s\}^\{p\}\_\{j\},\{d\(\\mathsf\{a\}^\{p\}\_\{j\},\\mathsf\{a\}^\{p\}\_\{i\}\)\},\\\|\\mathcal\{H\}^\{p\}\_\{j\}\-\\mathcal\{H\}^\{p\}\_\{i\}\\\|\_\{2\}\]\),whereNi:=\|๐\(๐บip\|๐ซ\)\|N\_\{i\}:=\|\\mbox\{$\\mathop\{\\mathbf\{N\}\}\\limits$\}\(\\mathsf\{a\}^\{p\}\_\{i\}\|\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\)\|, which denotes the number of neighboring pocket atoms of๐บip\\mathsf\{a\}^\{p\}\_\{i\}\. Thus,๐๐๐ฏ๐ฑ๐ซ\-โ๐ฝ๐พ๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{dec\}\}\\limitsestimates the positional noise of each atom๐บip\\mathsf\{a\}^\{p\}\_\{i\}by aggregating the relative position vectors from its neighboring atoms\. The weighting coefficientcjโic\_\{ji\}from each neighboring atom๐บjp\\mathsf\{a\}^\{p\}\_\{j\}is computed from the embeddings๐ฌp\\mathbf\{s\}^\{p\}andโp\\mathcal\{H\}^\{p\}produced by๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits\. This formulation exploits the local geometric structure of the pocket, forcing๐ฌp\\mathbf\{s\}^\{p\}andโp\\mathcal\{H\}^\{p\}to encode spatial relationships among neighboring atoms and reflect local geometric information\. We predict the added positional noise rather than the clean positions here since the noise follows a simpler and more tractable distribution, which stabilizes training and is easier to learn\.
For masked atom types,๐๐๐ฏ๐ฑ๐ซ\-โ๐ฝ๐พ๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{dec\}\}\\limitspredicts a probability vector over all possible atom types for each masked atom๐บip\\mathsf\{a\}^\{p\}\_\{i\}, denoted as๐ฏ^ip\\hat\{\\mathbf\{v\}\}^\{p\}\_\{i\}, by a multi\-layer perceptron \(MLP\):
๐ฏ^ip=softmaxโ\(MLPโ\(๐ฌip\)\)\.\\hat\{\\mathbf\{v\}\}^\{p\}\_\{i\}=\\text\{softmax\}\(\\text\{MLP\}\(\\mathbf\{s\}^\{p\}\_\{i\}\)\)\.\(9\)This formulation enables๐๐๐ฏ๐ฑ๐ซ\-โ๐ฝ๐พ๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{dec\}\}\\limitsto infer atom types based on the learned pocket embedding๐ฌp\\mathbf\{s\}^\{p\}\. By designing๐๐๐ฏ๐ฑ๐ซ\-โ๐ฝ๐พ๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{dec\}\}\\limitsto reconstruct atom types and positions, we guide๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limitsto learn pocket embeddings that capture both atomic characteristics and local geometry for accurate recovery and refinement\.
#### ๐๐๐ฏ๐ฑ๐ซ\\mathop\{\\mathsf\{msPRL\}\}\\limitsPretraining
To pretrain๐๐๐ฏ๐ฑ๐ซ\\mathop\{\\mathsf\{msPRL\}\}\\limits, followingUni\-Molโs denoising pretraining\[zhou2023uni\], we randomly corrupt 20% of the atoms in pocket๐ซ\\mathop\{\\mathcal\{P\}\}\\limits\. The original positions and atom types of the noisy pocket๐ซ\\mathop\{\\mathcal\{P\}\}\\limitsare reconstructed using๐๐๐ฏ๐ฑ๐ซ\-โ๐ฝ๐พ๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{dec\}\}\\limitsby minimizing the loss function:
โp:=1\|โณ\|โโ๐บipโ๐ซ๐โ\(๐บipโโณ\)โ\(โ๐ง^ipโ๐งipโ2\+๐งโ\(๐ฏ^ip,๐ฏip\)\),\\mathcal\{L\}\_\{p\}:=\\frac\{1\}\{\|\\mathcal\{M\}\|\}\\sum\_\{\\mathsf\{a\}^\{p\}\_\{i\}\\in\{\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\}\}\\mathbb\{I\}\(\\mathsf\{a\}^\{p\}\_\{i\}\\in\\mathcal\{M\}\)\\,\\Big\(\\,\\\|\\hat\{\\mathbf\{n\}\}^\{p\}\_\{i\}\-\\mathbf\{n\}^\{p\}\_\{i\}\\\|^\{2\}\\;\+\\;\\mathsf\{H\}\(\\hat\{\\mathbf\{v\}\}^\{p\}\_\{i\},\\mathbf\{v\}^\{p\}\_\{i\}\)\\Big\),\(10\)whereโณ\\mathcal\{M\}is the set of perturbed atoms in๐ซ\\mathop\{\\mathcal\{P\}\}\\limits,๐โ\(โ
\)\\mathbb\{I\}\(\\cdot\)denotes the indicator function,๐งโ\(โ
,โ
\)\\mathsf\{H\}\(\\cdot,\\cdot\)represents the cross\-entropy loss, and๐ง^ip\\hat\{\\mathbf\{n\}\}^\{p\}\_\{i\}and๐งip\\mathbf\{n\}^\{p\}\_\{i\}denote the predicted noise from the decoder and the ground\-truth noise from the corruption step, respectively\. The overall pretraining process is summarized in Algorithm[1](https://arxiv.org/html/2607.12349#alg1)\.
Algorithm 1๐๐๐ฏ๐ฑ๐ซ\\mathop\{\\mathsf\{msPRL\}\}\\limitsRequired Input: all๐ซ\\mathop\{\\mathcal\{P\}\}\\limitsin Training data of๐ข๐ฃ\\mathop\{\\mathsf\{CD\}\}\\limits
1:whilenot convergeddo
2:Sample a batch
โฌ\\mathcal\{B\}of๐ซ\\mathop\{\\mathcal\{P\}\}\\limits
3:
Lossโ0\\text\{Loss\}\\leftarrow 0
4:for
๐ซโโฌ\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\\in\\mathcal\{B\}do
5:
๐ฑp,๐ฏp,๐ฑr,๐ฏr\\mathbf\{x\}^\{p\},\\mathbf\{v\}^\{p\},\\mathbf\{x\}^\{r\},\\mathbf\{v\}^\{r\}=
addNoiseโ\(๐ซ\)\\text\{addNoise\}\(\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\)โณ\\trianglerightperturb pocket structure
6:
๐ฌp,โp=๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\(๐ฑp,๐ฏp,๐ฑr,๐ฏr\)\\mathbf\{s\}^\{p\},\\mathcal\{H\}^\{p\}=\\mbox\{$\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits$\}\(\\mathbf\{x\}^\{p\},\\mathbf\{v\}^\{p\},\\mathbf\{x\}^\{r\},\\mathbf\{v\}^\{r\}\)โณ\\trianglerightencode the positions and features into latent embeddings
7:
๐ซ^=๐๐๐ฏ๐ฑ๐ซ\-โ๐ฝ๐พ๐ผ\(๐ฌp,โp,๐ฑp\)\\hat\{\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\}=\\mbox\{$\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{dec\}\}\\limits$\}\(\\mathbf\{s\}^\{p\},\\mathcal\{H\}^\{p\},\\mathbf\{x\}^\{p\}\)โณ\\trianglerightreconstruct pocket by decoding the embeddings
8:
LossโLoss\+โpโ\(๐ซ^,๐ซ\)\\text\{Loss\}\\leftarrow\\text\{Loss\}\+\\mathcal\{L\}\_\{p\}\(\\hat\{\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\},\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\)โณ\\trianglerightaccumulate reconstruction loss \(Equation[10](https://arxiv.org/html/2607.12349#Sx6.E10)\)
9:endfor
10:
Lossโ1\|โฌ\|โLoss\\text\{Loss\}\\leftarrow\\frac\{1\}\{\|\\mathcal\{B\}\|\}\\text\{Loss\}
11:
Loss\.backward\(๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ,๐๐๐ฏ๐ฑ๐ซ\-โ๐ฝ๐พ๐ผ\)\\text\{Loss\}\.\\text\{backward\(\\mbox\{$\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits$\},\\;\\mbox\{$\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{dec\}\}\\limits$\}\)\}โณ\\trianglerightEquation[10](https://arxiv.org/html/2607.12349#Sx6.E10)
12:Update๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits,๐๐๐ฏ๐ฑ๐ซ\-โ๐ฝ๐พ๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{dec\}\}\\limits
13:endwhile
14:return๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits
### Pocket\-conditioned Ligand Generation via Conditional Diffusion \(๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limits\)
We present a new conditional diffusion model,๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limits, for the generation of three\-dimensional \(3D\) ligands that bind to specific, given protein pockets, referred to as condition pockets\. The proposed model generates realistic 3D ligands with high binding affinities by:\(1\)incorporating learned pocket structures and interactions within the pocket from๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits,\(2\)modeling atomic interactions within the ligand, and\(3\)learning pocket\-ligand interactions\. The subsequent sections provide a detailed description of the diffusion process \(Section Diffusion Process\), the training methodology of๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limits\(Section Model Training\), and the pocket\-conditioned ligand generator \(Section Pocket\-conditioned Ligand Generator \(๐๐ผ๐ซ๐ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limits\)\)\.
Following the framework of denoising diffusion probabilistic models\[ho2020ddpm\],๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitscomprises a forward diffusion process that progressively adds noise to the atomic positions\{๐ฑig\}\\\{\\mathbf\{x\}^\{g\}\_\{i\}\\\}and features\{๐ฏig\}\\\{\\mathbf\{v\}^\{g\}\_\{i\}\\\}of ligands, and a reverse generative process that is trained to denoise and generate new ligands during inference\. During training, the model learns to invert the forward corruption process by iteratively denoising noisy ligands\. During inference,๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsfirst samples noisy ligand atom positions and features at stepTTfrom predefined simple distributions and then reconstructs realistic 3D ligand structures by iteratively removing noise untilt=0t=0\. At each reverse steptt, given a noisy structure\{\(๐ฑi,tg,๐ฏi,tg\)\}\\\{\(\\mathbf\{x\}^\{g\}\_\{i,t\},\\mathbf\{v\}^\{g\}\_\{i,t\}\)\\\},๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsemploys a GNN\-based ligand generator that uses fixed pocket embeddings from๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limitstogether with pocket atom/residue positions and types to predict the noise\-free atom positions\{๐ฑi,0g\}\\\{\\mathbf\{x\}^\{g\}\_\{i,0\}\\\}and features\{๐ฏi,0g\}\\\{\\mathbf\{v\}^\{g\}\_\{i,0\}\\\}of the ligand atoms\. Unlike other diffusion\-based SBDD methods\[guan2023targetdiff,guan2023decompdiff,huang2024protein,gu2024aligning\],๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limits\(1\)leverages pocket positions and types exclusively to model pocket\-ligand interactions during denoising, while the pocket structure remains captured in fixed embeddings produced by๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits\. The fixed pocket representation prevents embedding drifts and re\-encoding artifacts across๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsโs denoising process, so the ligand generator can learn pocket\-ligand interactions efficiently and with greater coherence across the generation process\. Like prior SBDD works\[guan2023targetdiff,guan2023decompdiff,huang2024protein,gu2024aligning\],๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limits\(2\)models the atomic interactions between the ligand and pocket\. In addition to atom\-level modeling,๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsfurther integrates residue\-level interactions\. This multi\-scale featured design provides the ligand with a hierarchical view of the pocket, enabling a comprehensive understanding of pocket\-ligand interactions\.
#### Forward Diffusion Process \(๐ผ๐๐๐ฃ๐๐๐บ๐โ\-โ๐ฟ๐๐๐๐บ๐๐ฝ\\mathop\{\\mathsf\{conDitar\}\\text\{\-\}\\mathsf\{forward\}\}\\limits\)
Following the previous work\[guan2023targetdiff\], each ligand atom position is progressively noised according to a Gaussian transition in the forward process\[ho2020ddpm\]:
qโ\(๐ฑi,tgโฃ๐ฑi,tโ1g\)=๐ฉโ\(๐ฑi,tg;1โฮฒt๐ฑโ๐ฑi,tโ1g,ฮฒt๐ฑโ๐\),q\(\\mathbf\{x\}^\{g\}\_\{i,t\}\\mid\\mathbf\{x\}^\{g\}\_\{i,t\-1\}\)=\\mathcal\{N\}\\\!\\left\(\\mathbf\{x\}^\{g\}\_\{i,t\};\\sqrt\{1\-\\beta^\{\\mathbf\{x\}\}\_\{t\}\}\\,\\mathbf\{x\}^\{g\}\_\{i,t\-1\},\\beta^\{\\mathbf\{x\}\}\_\{t\}\\mathbf\{I\}\\right\),\(11\)whereฮฒt๐ฑ\\beta^\{\\mathbf\{x\}\}\_\{t\}is a predefined variance schedule that controls the noise magnitude added at steptt;๐\\mathbf\{I\}denotes the identity matrix\. OverTTsteps, this produces a Markov chain that gradually transforms the clean coordinates into isotropic Gaussian noise\. Similarly, atom types are corrupted through a categorical diffusion process\. At each steptt, the current atom type is retained with probability1โฮฒt๐ฏ1\-\\beta^\{\\mathbf\{v\}\}\_\{t\}, while the remaining probability mass is uniformly distributed among the other possible types:
qโ\(๐ฏi,tgโฃ๐ฏi,tโ1g\)=๐โ\(๐ฏi,tg;\(1โฮฒt๐ฏ\)โ๐ฏi,tโ1g\+ฮฒt๐ฏdaโ๐\),q\(\\mathbf\{v\}^\{g\}\_\{i,t\}\\mid\\mathbf\{v\}^\{g\}\_\{i,t\-1\}\)=\\mathcal\{C\}\\\!\\Big\(\\mathbf\{v\}^\{g\}\_\{i,t\};\\,\(1\-\\beta^\{\\mathbf\{v\}\}\_\{t\}\)\\,\{\\mathbf\{v\}^\{g\}\_\{i,t\-1\}\}\\;\+\\;\\frac\{\\beta^\{\\mathbf\{v\}\}\_\{t\}\}\{d\_\{a\}\}\\mathbf\{1\}\\Big\),\(12\)whereฮฒt๐ฏ\\beta^\{\\mathbf\{v\}\}\_\{t\}controls the level of corruption;๐\\mathbf\{1\}denotes the all\-ones vector indicating uniform probability over alldad\_\{a\}atom types\.
#### Reverse Generative Process \(๐ผ๐๐๐ฃ๐๐๐บ๐โ\-โ๐๐พ๐๐พ๐๐๐พ\\mathop\{\\mathsf\{conDitar\}\\text\{\-\}\\mathsf\{reverse\}\}\\limits\)
In the reverse process,๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitslearns to reverse the forward process and generates ligands by denoising from\{\(๐ฑi,tg,๐ฏi,tg\)\}\\\{\(\\mathbf\{x\}^\{g\}\_\{i,t\},\\mathbf\{v\}^\{g\}\_\{i,t\}\)\\\}to\{\(๐ฑi,tโ1g,๐ฏi,tโ1g\)\}\\\{\(\\mathbf\{x\}^\{g\}\_\{i,t\-1\},\\mathbf\{v\}^\{g\}\_\{i,t\-1\}\)\\\}\. Following Ho*et al\.*\[ho2020ddpm\],๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsapproximates the probability of \(๐ฑi,tโ1g\\mathbf\{x\}^\{g\}\_\{i,t\-1\},๐ฏi,tโ1g\\mathbf\{v\}^\{g\}\_\{i,t\-1\}\) denoised from \(๐ฑi,tg\\mathbf\{x\}^\{g\}\_\{i,t\},๐ฏi,tg\\mathbf\{v\}^\{g\}\_\{i,t\}\) using the posteriorpโ\(๐ฑi,tโ1g\|๐ฑi,tg,๐ฑi,0g\)p\(\\mathbf\{x\}^\{g\}\_\{i,t\-1\}\|\\mathbf\{x\}^\{g\}\_\{i,t\},\\mathbf\{x\}^\{g\}\_\{i,0\}\)andpโ\(๐ฏi,tโ1g\|๐ฏi,tg,๐ฏi,0g\)p\(\\mathbf\{v\}^\{g\}\_\{i,t\-1\}\|\\mathbf\{v\}^\{g\}\_\{i,t\},\\mathbf\{v\}^\{g\}\_\{i,0\}\)\. Given that๐ฑi,0g\\mathbf\{x\}^\{g\}\_\{i,0\}and๐ฏi,0g\\mathbf\{v\}^\{g\}\_\{i,0\}are unknown in the backward process, at each steptt,๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsuses a new pocket\-conditioned ligand generator,๐๐ผ๐ซ๐ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limits\(Pocket\-conditioned Ligand Generator \(๐๐ผ๐ซ๐ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limits\) section\), to estimate them as follows:
\{\(๐ฑ~i,0,tg,๐ฏ~i,0,tg\)\}=๐๐ผ๐ซ๐ฆ\(\{๐ฑi,tg\},\{๐ฏi,tg\},\{๐ฌjp\},\{โjp\}\),\\\{\(\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{i,0,t\},\\tilde\{\\mathbf\{v\}\}^\{g\}\_\{i,0,t\}\)\\\}=\\mbox\{$\\mathop\{\\mathsf\{pcLG\}\}\\limits$\}\(\\\{\\mathbf\{x\}^\{g\}\_\{i,t\}\\\},\\\{\\mathbf\{v\}^\{g\}\_\{i,t\}\\\},\\\{\\mathbf\{s\}^\{p\}\_\{j\}\\\},\\\{\\mathcal\{H\}^\{p\}\_\{j\}\\\}\),\(13\)where๐ฑ~i,0,tg\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{i,0,t\}and๐ฏ~i,0,tg\\tilde\{\\mathbf\{v\}\}^\{g\}\_\{i,0,t\}are the estimates of๐ฑi,0g\\mathbf\{x\}^\{g\}\_\{i,0\}and๐ฏi,0g\\mathbf\{v\}^\{g\}\_\{i,0\}at timesteptt, respectively\.\{๐ฌjp\}\\\{\\mathbf\{s\}^\{p\}\_\{j\}\\\}and\{โjp\}\\\{\\mathcal\{H\}^\{p\}\_\{j\}\\\}represent the pretrained scalar and vector pocket embeddings\)\. Using the estimates,๐ฑi,tโ1g\\mathbf\{x\}^\{g\}\_\{i,t\-1\}can be sampled as:
pโ\(๐ฑi,tโ1gโฃ๐ฑi,tg\)\\displaystyle p\(\\mathbf\{x\}^\{g\}\_\{i,t\-1\}\\mid\\mathbf\{x\}^\{g\}\_\{i,t\}\)โqโ\(๐ฑi,tโ1gโฃ๐ฑi,tg,๐ฑ~i,0,tg\)=๐ฉโ\(๐ฑi,tโ1gโฃฮผโ\(๐ฑi,tg,๐ฑ~i,0,tg\),ฮฒ~t๐ฑโ๐\),\\displaystyle\\approx q\(\\mathbf\{x\}^\{g\}\_\{i,t\-1\}\\mid\\mathbf\{x\}^\{g\}\_\{i,t\},\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{i,0,t\}\)=\\mathcal\{N\}\(\\mathbf\{x\}^\{g\}\_\{i,t\-1\}\\mid\\mu\(\\mathbf\{x\}^\{g\}\_\{i,t\},\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{i,0,t\}\),\\tilde\{\\beta\}\_\{t\}^\{\\mathbf\{x\}\}\\mathbf\{I\}\),\(14\)whereฮผโ\(๐ฑi,tg,๐ฑ~i,0,tg\)\\mu\(\\mathbf\{x\}^\{g\}\_\{i,t\},\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{i,0,t\}\)is the estimated mean of Gaussian transition distribution \(detailed in Appendix[C](https://arxiv.org/html/2607.12349#A3)\)\. Similarly,๐ฏi,tโ1g\\mathbf\{v\}^\{g\}\_\{i,t\-1\}can be sampled as
pโ\(๐ฏi,tโ1gโฃ๐ฏi,tg\)โqโ\(๐ฏi,tโ1gโฃ๐ฏi,tg,๐ฏ~i,0,tg\)=๐โ\(๐ฏi,tโ1gโฃ๐โ\(๐ฏi,tg,๐ฏ~i,0,tg\)\),p\(\\mathbf\{v\}^\{g\}\_\{i,t\-1\}\\mid\\mathbf\{v\}^\{g\}\_\{i,t\}\)\\approx q\(\\mathbf\{v\}^\{g\}\_\{i,t\-1\}\\mid\\mathbf\{v\}^\{g\}\_\{i,t\},\\tilde\{\\mathbf\{v\}\}^\{g\}\_\{i,0,t\}\)=\\mathcal\{C\}\(\\mathbf\{v\}^\{g\}\_\{i,t\-1\}\\mid\\mathbf\{c\}\(\\mathbf\{v\}^\{g\}\_\{i,t\},\\tilde\{\\mathbf\{v\}\}^\{g\}\_\{i,0,t\}\)\),\(15\)where๐โ\(๐ฏi,tg,๐ฏ~i,0,tg\)\\mathbf\{c\}\\big\(\\mathbf\{v\}^\{g\}\_\{i,t\},\\tilde\{\\mathbf\{v\}\}^\{g\}\_\{i,0,t\}\\big\)denotes the estimated probability vector that parameterizes the categorical transition distribution \(detailed in Appendix[C](https://arxiv.org/html/2607.12349#A3)\)\. In practice, the denoising update for๐ฑi,tโ1g\\mathbf\{x\}^\{g\}\_\{i,t\-1\}can be calculated as
๐ฑi,tโ1g=ฮผโ\(๐ฑi,tg,๐ฑ~i,0,tg\)\+ฮฒ~t๐ฑโ๐ณi,tโ1๐ฑ,\\mathbf\{x\}^\{g\}\_\{i,t\-1\}=\\mu\(\\mathbf\{x\}^\{g\}\_\{i,t\},\\tilde\{\\mathbf\{x\}\}\_\{i,0,t\}^\{g\}\)\+\\sqrt\{\\tilde\{\\beta\}^\{\\mathbf\{x\}\}\_\{t\}\}\\mathbf\{z\}^\{\\mathbf\{x\}\}\_\{i,t\-1\},\(16\)where๐ณi,tโ1๐ฑโผ๐ฉโ\(๐,๐\)\\mathbf\{z\}^\{\\mathbf\{x\}\}\_\{i,t\-1\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)is a standard Gaussian noise sampled at timett, and the denoising update for๐ฏi,tโ1g\\mathbf\{v\}^\{g\}\_\{i,t\-1\}can be calculated as
๐ฏi,tโ1g=onehotโก\(argโกmaxโโก\[logโก๐โ\(๐ฏi,tg,๐ฏ~i,0,tg\)โ\+๐ณi,tโ1,โ๐\]\),\\mathbf\{v\}^\{g\}\_\{i,t\-1\}=\\operatorname\{onehot\}\\\!\\left\(\\arg\\max\_\{\\ell\}\\big\[\\,\\log\\mathbf\{c\}\(\\mathbf\{v\}^\{g\}\_\{i,t\},\\tilde\{\\mathbf\{v\}\}^\{g\}\_\{i,0,t\}\)\_\{\\ell\}\+\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{i,t\-1,\\ell\}\\,\\big\]\\right\),\(17\)where๐ณi,tโ1,โ๐ฏโผGumbelโ\(0,1\)\\mathbf\{z\}\_\{i,t\-1,\\ell\}^\{\\mathbf\{v\}\}\\sim\\text\{Gumbel\}\(0,1\)forโ=1,โฆ,da\\ell=1,\\dots,d\_\{a\}\(dad\_\{a\}is the total number of atom types\) are independent Gumbel noises\. Here, we use the Gumbel\-max trick to sample from the discrete categorical distribution\. More details of the backward process are available in Appendix[C](https://arxiv.org/html/2607.12349#A3)\.
##### Determining Ligand Sizes
During the generative process, the ligand size \(i\.e\., the number of heavy atoms\) is unknown and must be determined before atom positions and types can be generated\. The ligand size is sampled based on the volume of the condition pocket, referred to as the pocket size\. The pocket size is determined by the region in proximity to the reference ligand\. Pocket atoms that are among the three nearest neighbors of the ligand atoms are selected, and the maximum pairwise distance among them defines the pocket size\. Pocket sizes are grouped into ordered intervals, each associated with an empirical distribution of ligand atom counts from the training data\. The ligand size is then sampled from the empirical distribution corresponding to the pocketโs interval\.
#### ๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsModel Training
๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsis trained to recover original ligand atom positions๐ฑi,0g\\mathbf\{x\}^\{g\}\_\{i,0\}and features๐ฏi,0g\\mathbf\{v\}^\{g\}\_\{i,0\}, conditioned on the protein pocket\. Specifically,๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsuses the combination of the following three losses as its loss function\.
##### Atom Position Loss
This loss measures the Euclidean errors between the predicted positions๐ฑ~i,0,tg\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{i,0,t\}at stepttand the ground\-truth positions๐ฑi,0g\\mathbf\{x\}^\{g\}\_\{i,0\}as follows:
โt๐ฑโ\(๐\)\\displaystyle\\mathcal\{L\}\_\{t\}^\{\\mathbf\{x\}\}\(\)=wt๐ฑโโ๐บigโ๐KL\(q\(๐ฑi,tโ1g\|๐ฑi,tg,๐ฑi,0g\)\|\|q\(๐ฑi,tโ1g\|๐ฑi,tg,๐ฑ~i,0,tg\)\)\\displaystyle=w\_\{t\}^\{\\mathbf\{x\}\}\\sum\_\{\\forall\\mathsf\{a\}^\{g\}\_\{i\}\\in\{\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\}\}\\text\{KL\}\(q\(\\mathbf\{x\}^\{g\}\_\{i,t\-1\}\|\\mathbf\{x\}^\{g\}\_\{i,t\},\\mathbf\{x\}^\{g\}\_\{i,0\}\)\|\|q\(\\mathbf\{x\}^\{g\}\_\{i,t\-1\}\|\\mathbf\{x\}^\{g\}\_\{i,t\},\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{i,0,t\}\)\)\(18\)=wt๐ฑโโโ๐บigโ๐โ๐ฑ~i,0,tgโ๐ฑi,0gโ,\\displaystyle=w\_\{t\}^\{\\mathbf\{x\}\}\\sum\_\{\\forall\\mathsf\{a\}^\{g\}\_\{i\}\\in\{\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\}\}\\\|\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{i,0,t\}\-\\mathbf\{x\}^\{g\}\_\{i,0\}\\\|,wherewt๐ฑw\_\{t\}^\{\\mathbf\{x\}\}is a time\-dependent weight at steptt, calculated by the signal\-to\-noise ratio\[kingma2021variational\]\(detailed in Appendix[B](https://arxiv.org/html/2607.12349#A2)\), and KL is the Kullback\-Leibler divergence\[kullback1951information\]\. Asttincreases,ฮฑยฏt๐ฑ\\bar\{\\alpha\}\_\{t\}^\{\\mathbf\{x\}\}decreases monotonically, resulting in an increase in noise level and a corresponding decrease inwt๐ฑw\_\{t\}^\{\\mathbf\{x\}\}until it reaches the thresholdฮด\\delta\. As a result, it encourages the model to emphasize more accurate reconstruction of molecular structures when the data have low noise levels and contain sufficient signals\.
##### Atom Type Loss
This loss measures the KL divergence\[kullback1951information\]between the ground\-truth posteriorqโ\(๐ฏi,tโ1g\|๐ฏi,tg,๐ฏi,0g\)q\(\\mathbf\{v\}^\{g\}\_\{i,t\-1\}\|\\mathbf\{v\}^\{g\}\_\{i,t\},\\mathbf\{v\}^\{g\}\_\{i,0\}\)and its estimateqโ\(๐ฏi,tโ1g\|๐ฏtg,๐ฏ~i,0,tg\)q\(\\mathbf\{v\}^\{g\}\_\{i,t\-1\}\|\\mathbf\{v\}^\{g\}\_\{t\},\\tilde\{\\mathbf\{v\}\}^\{g\}\_\{i,0,t\}\)as follows:
โt๐ฏโ\(๐\)\\displaystyle\\mathcal\{L\}\_\{t\}^\{\\mathbf\{v\}\}\(\)=โโ๐บigโ๐KL\(q\(๐ฏi,tโ1g\|๐ฏi,tg,๐ฏi,0g\)\|\|q\(๐ฏi,tโ1g\|๐ฏi,tg,๐ฏ~i,0,tg\)\)\\displaystyle=\\sum\_\{\\forall\\mathsf\{a\}^\{g\}\_\{i\}\\in\{\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\}\}\\text\{KL\}\(q\(\\mathbf\{v\}^\{g\}\_\{i,t\-1\}\|\\mathbf\{v\}^\{g\}\_\{i,t\},\\mathbf\{v\}^\{g\}\_\{i,0\}\)\|\|q\(\\mathbf\{v\}^\{g\}\_\{i,t\-1\}\|\\mathbf\{v\}^\{g\}\_\{i,t\},\\tilde\{\\mathbf\{v\}\}^\{g\}\_\{i,0,t\}\)\)\(19\)=โโ๐บigโ๐KL\(๐\(๐ฏi,tg,๐ฏi,0g\)\|\|๐\(๐ฏi,tg,๐ฏ~i,0,tg\)\),\\displaystyle=\\sum\_\{\\forall\\mathsf\{a\}^\{g\}\_\{i\}\\in\{\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\}\}\\text\{KL\}\(\\mathbf\{c\}\(\\mathbf\{v\}^\{g\}\_\{i,t\},\\mathbf\{v\}^\{g\}\_\{i,0\}\)\|\|\\mathbf\{c\}\(\\mathbf\{v\}^\{g\}\_\{i,t\},\\tilde\{\\mathbf\{v\}\}^\{g\}\_\{i,0,t\}\)\),where๐โ\(๐ฏi,tg,๐ฏi,0g\)\\mathbf\{c\}\(\\mathbf\{v\}^\{g\}\_\{i,t\},\\mathbf\{v\}^\{g\}\_\{i,0\}\)is the categorical distribution of๐ฏi,tg\\mathbf\{v\}^\{g\}\_\{i,t\}and๐ฮโ\(๐ฏi,tg,๐ฏ~i,0,tg\)\\mathbf\{c\}\_\{\\Theta\}\(\\mathbf\{v\}^\{g\}\_\{i,t\},\\tilde\{\\mathbf\{v\}\}^\{g\}\_\{i,0,t\}\)is an estimate of๐โ\(๐ฏi,tg,๐ฏi,0g\)\\mathbf\{c\}\(\\mathbf\{v\}^\{g\}\_\{i,t\},\\mathbf\{v\}^\{g\}\_\{i,0\}\)\.
##### Bond Type Loss
Following the literature\[chen2025generating\], this loss is to help๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsbetter understand the relations among atoms\. By explicitly learning bonds, the model better enforces chemical validity and yields more realistic molecular structures\. To achieve this, at each time steptt, for eachll\-th layer of ligand generator๐๐ผ๐ซ๐ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limits\(Pocket\-conditioned Ligand Generator \(๐๐ผ๐ซ๐ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limits\) section\),๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsis optimized to accurately predict the bond types among pairs of ligand atoms\. Specifically, the bond type loss is defined as follows:
โt,l๐โ\(๐\)=โโ๐บigโ๐โโ๐บjgโ๐\(๐บig\|๐\)๐งโ\(๐iโj,t,l,๐iโj\),\\mathcal\{L\}^\{\\mathtt\{b\}\}\_\{t,l\}\(\{\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\}\)=\\sum\_\{\\forall\\mathsf\{a\}^\{g\}\_\{i\}\\in\{\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\}\}~\\sum\_\{\\forall\\mathsf\{a\}^\{g\}\_\{j\}\\in\{\\mbox\{$\\mathop\{\\mathbf\{N\}\}\\limits$\}\(\\mathsf\{a\}^\{g\}\_\{i\}\|\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\)\}\}\\mathsf\{H\}\(\\mbox\{$\\mathop\{\\mathbf\{e\}\}\\limits$\}\_\{ij,t,l\},\\mbox\{$\\mathop\{\\mathbf\{b\}\}\\limits$\}\_\{ij\}\),\(20\)where๐งโ\(โ
\)\\mathsf\{H\}\(\\cdot\)denotescross\-entropyloss,๐iโj,t,l\\mbox\{$\\mathop\{\\mathbf\{e\}\}\\limits$\}\_\{ij,t,l\}represents the bond type predictions \(Equation[27](https://arxiv.org/html/2607.12349#Sx6.E27)\) between atoms๐บig\\mathsf\{a\}^\{g\}\_\{i\}and๐บjg\\mathsf\{a\}^\{g\}\_\{j\}at thell\-th layer at time steptt;๐\(๐บig\|๐\)\\mbox\{$\\mathop\{\\mathbf\{N\}\}\\limits$\}\(\\mathsf\{a\}^\{g\}\_\{i\}\|\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\)denotes thekk\-nearest ligand atoms of atom๐บig\\mathsf\{a\}^\{g\}\_\{i\}in position๐ฑi,tg\\mathbf\{x\}^\{g\}\_\{i,t\}; and๐iโj\\mbox\{$\\mathop\{\\mathbf\{b\}\}\\limits$\}\_\{ij\}denotes one\-hot vector indicating the ground\-truth bond type between๐บig\\mathsf\{a\}^\{g\}\_\{i\}and๐บjg\\mathsf\{a\}^\{g\}\_\{j\}\. The total bond type prediction loss is computed by aggregating losses across different layers as follows:
โt๐โ\(๐\)=wt๐ฑLโ1โโl=1Lโ1โt,l๐โ\(๐\)\+wt๐ฑโโt,L๐โ\(๐\),\\mathcal\{L\}^\{\\mathbf\{b\}\}\_\{t\}\(\{\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\}\)=\\frac\{w\_\{t\}^\{\\mathbf\{x\}\}\}\{L\-1\}\\sum\_\{l=1\}^\{L\-1\}\\mathcal\{L\}^\{\\mathbf\{b\}\}\_\{t,l\}\(\{\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\}\)\+w\_\{t\}^\{\\mathbf\{x\}\}\\mathcal\{L\}^\{\\mathbf\{b\}\}\_\{t,L\}\(\{\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\}\),\(21\)wherewt๐ฑw\_\{t\}^\{\\mathbf\{x\}\}is the same timestep weight used in the atom position loss;LLrepresents the total number of layers in thepocket\-conditionedligand generator๐๐ผ๐ซ๐ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limits\. Similar toโt๐ฑโ\(๐\)\\mathcal\{L\}^\{\\mathbf\{x\}\}\_\{t\}\(\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\)\(Equation[18](https://arxiv.org/html/2607.12349#Sx6.E18)\), the weightwt๐ฑw\_\{t\}^\{\\mathbf\{x\}\}encourages the model to focus more on accurately predicting bond types when the data provides sufficient signals, rather than being confused by major noises in the data\. Note that, similar to Jumper*et al\.*\[Jumper2021\], in Equation[21](https://arxiv.org/html/2607.12349#Sx6.E21),๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsassigns different weights to the final layer \(i\.e\.,ll=LL\) compared to the other layers, as we empirically find this design benefits the generation performance\.
##### Overall๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsloss
The overall loss function for๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsis defined as follows:
โdiff=๐ผ๐โผ๐ณโ๐ผtโผ๐ฐโ\(1,T\)โ\(โt๐ฑโ\(๐\)\+ฮพโโt๐ฏโ\(๐\)\+ฮถโโt๐โ\(๐\)\),\\mathcal\{L\}\_\{\\text\{diff\}\}=\\mathbb\{E\}\_\{\{\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\\sim\\text\{$\\mathtt\{D\}$\}\}\}\\mathbb\{E\}\_\{t\\sim\\mathcal\{U\}\(1,T\)\}\\big\(\\mathcal\{L\}^\{\\mathbf\{x\}\}\_\{t\}\(\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\)\+\\xi\\,\\mathcal\{L\}^\{\\mathbf\{v\}\}\_\{t\}\(\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\)\+\\zeta\\,\\mathcal\{L\}^\{\\mathbf\{b\}\}\_\{t\}\(\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\)\\big\),\(22\)where๐ณ\\mathtt\{D\}is the training binding complexes;T\{T\}is the total number of diffusion timesteps;๐ฐโ\(1,T\)\\mathcal\{U\}\(1,T\)represents uniform distribution over these timesteps;ฮพ\>0\\xi\>0andฮถ\>0\\zeta\>0are twohyper\-parametersthat balanceโt๐ฑ\\mathcal\{L\}^\{\\mathbf\{x\}\}\_\{t\}\(๐\\mathop\{\\mathcal\{D\}\}\\limits\),โt๐ฏ\\mathcal\{L\}^\{\\mathbf\{v\}\}\_\{t\}\(๐\\mathop\{\\mathcal\{D\}\}\\limits\) andโt๐โ\(๐\)\\mathcal\{L\}^\{\\mathbf\{b\}\}\_\{t\}\(\{\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\}\)\.
### Pocket\-conditioned Ligand Generator \(๐๐ผ๐ซ๐ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limits\)
๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsemploys a pocket\-conditioned ligand generator,๐๐ผ๐ซ๐ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limits, to denoise noisy ligand atom positions๐ฑ~i,tg\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{i,t\}and features๐ฏ~i,tg\\tilde\{\\mathbf\{v\}\}^\{g\}\_\{i,t\}at each diffusion step\. At a high level, denoising involves predicting the added noises and recovering the original ligand structures by leveraging chemical information and spatial information, including pocket geometry encoded by pretrained๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits, ligand geometry, and the ligandโs relative spatial arrangement within the pocket\. To guide this process,๐๐ผ๐ซ๐ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitslearns expressive ligand atom representations by modeling two types of interactions:\(1\)interactions within the ligand, which model relationships among neighboring atoms to capture each atomโs local environment, and\(2\)pocket\-ligand interactions, which capture pocket circumstances to provide pocket\-informed ligand representations\. Through modeling these interactions,๐๐ผ๐ซ๐ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitscaptures both the ligandโs structure and its external interaction patterns with the pocket\. Furthermore, by pretraining pocket representation,๐๐ผ๐ซ๐ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitsavoids modeling pocket structure and instead focuses on learning intra\-ligand and pocket\-ligand interactions\. When no ambiguity arises, we will eliminate subscriptttin the notations and use\(๐ฑig,๐ฏig\)\(\\mathbf\{x\}^\{g\}\_\{i\},\\mathbf\{v\}^\{g\}\_\{i\}\)for brevity\.
#### Ligand Representation Learning
This section describes how๐๐ผ๐ซ๐ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitslearns ligand representations \(embeddings with superscriptg\) that capture both internal molecular structure and interactions with the pocket\.๐๐ผ๐ซ๐ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitsemploys a multi\-layer GNN to iteratively update ligand atom embeddings based on intra\-ligand and pocket\-ligand interactions\. Specifically, at thell\-th layer of the GNN, for each ligand atom๐บig\\mathsf\{a\}^\{g\}\_\{i\},๐๐ผ๐ซ๐ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitslearns an invariant scalar embedding๐ฌi,lgโโdร1\\mathbf\{s\}^\{g\}\_\{i,l\}\\in\\mathbb\{R\}^\{d\\times 1\}to capture atom features and an equivariant vector embeddingโi,lgโโdร3\\mathcal\{H\}^\{g\}\_\{i,l\}\\in\\mathbb\{R\}^\{d\\times 3\}for 3D structural features\. For๐บig\\mathsf\{a\}^\{g\}\_\{i\},๐๐ผ๐ซ๐ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitsupdates๐ฌi,lg\\mathbf\{s\}^\{g\}\_\{i,l\}andโi,lg\\mathcal\{H\}^\{g\}\_\{i,l\}by aggregating information from neighboring atoms of๐บig\\mathsf\{a\}^\{g\}\_\{i\}in the current noisy ligand and incorporating interactions with neighboring pocket atoms\. The embeddings๐ฌi,lg\\mathbf\{s\}^\{g\}\_\{i,l\}andโi,lg\\mathcal\{H\}^\{g\}\_\{i,l\}are updated as follows:
\(๐ฌi,lg,โi,lg\)=GVPโ\(๐กi,l,๐ดi,l\),where\(\\mathbf\{s\}^\{g\}\_\{i,l\},\\mathcal\{H\}^\{g\}\_\{i,l\}\)=\\text\{GVP\}\(\\mathbf\{h\}\_\{i,l\},\\mathcal\{Y\}\_\{i,l\}\),\\text\{ where \}\(23\)๐กi,l\\displaystyle\\mathbf\{h\}\_\{i,l\}=\[๐ฏig,๐ฌi,lโ1gโfrom the ligand,๐ฌi,lโ1c,๐ฌi,lโ1a,โfrom the pocketโ๐บjgโ๐\(๐บig\|๐\)zjโi,l๐ฆjโi,lg\],\\displaystyle=\\biggr\[\\underbrace\{\\mathbf\{v\}^\{g\}\_\{i\},\\mathbf\{s\}^\{g\}\_\{i,l\-1\}\}\_\{\\mathclap\{\\text\{from the ligand\}\}\},~~~\\underbrace\{\\mathbf\{s\}^\{c\}\_\{i,l\-1\},\\mathbf\{s\}^\{a\}\_\{i,l\-1\},\}\_\{\\mathclap\{\\text\{from the pocket\}\}\}~\\sum\_\{\\mathsf\{a\}^\{g\}\_\{j\}\\in\{\\mbox\{$\\mathop\{\\mathbf\{N\}\}\\limits$\}\(\\mathsf\{a\}^\{g\}\_\{i\}\|\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\)\}\}z\_\{ji,l\}\\mathbf\{m\}^\{g\}\_\{ji,l\}\\biggr\],\(24\)๐ดi,l\\displaystyle\\mathcal\{Y\}\_\{i,l\}=\[๐ฑig,โi,lโ1gโfrom the ligand,โi,lโ1c,โi,lโ1a,โfrom the pocketโ๐บjgโ๐\(๐บig\|๐\)zjโi,lMjโi,lg\],\\displaystyle=\\biggr\[\\underbrace\{\\mathbf\{x\}^\{g\}\_\{i\},\\mathcal\{H\}^\{g\}\_\{i,l\-1\}\}\_\{\\mathclap\{\\text\{from the ligand\}\}\},~\\underbrace\{\\mathcal\{H\}^\{c\}\_\{i,l\-1\},\\mathcal\{H\}^\{a\}\_\{i,l\-1\},\}\_\{\\mathclap\{\\text\{from the pocket\}\}\}\\sum\_\{\\mathsf\{a\}^\{g\}\_\{j\}\\in\{\\mbox\{$\\mathop\{\\mathbf\{N\}\}\\limits$\}\(\\mathsf\{a\}^\{g\}\_\{i\}\|\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\)\}\}z\_\{ji,l\}M^\{g\}\_\{ji,l\}\\biggr\],wherecrepresents embeddings that model pocket\-ligand atomic interactions, andarepresents residue\-level interaction embeddings\.๐ฌi,lโ1c\\mathbf\{s\}^\{c\}\_\{i,l\-1\}andโi,lโ1c\\mathcal\{H\}^\{c\}\_\{i,l\-1\}denote the scalar and vector embeddings that encode how ligand atom๐บig\\mathsf\{a\}^\{g\}\_\{i\}interacts with neighboring pocket atoms from the\(lโ1\)\(l\-1\)\-th layer of the GNN \(detailed in Equation[29](https://arxiv.org/html/2607.12349#Sx6.E29)\);๐ฌi,lโ1a\\mathbf\{s\}^\{a\}\_\{i,l\-1\}andโi,lโ1a\\mathcal\{H\}^\{a\}\_\{i,l\-1\}denote the scalar and vector embeddings that capture structures from surrounding pocket residues for ligand atom๐บig\\mathsf\{a\}^\{g\}\_\{i\}\(detailed in Equation[30](https://arxiv.org/html/2607.12349#Sx6.E30)\);๐ฆjโi,lg\\mathbf\{m\}^\{g\}\_\{ji,l\}andMjโi,lgM^\{g\}\_\{ji,l\}represent the scalar and vector message embeddings in the\(lโ1\)\(l\-1\)\-th layer of the GNN to propagate information from๐บig\\mathsf\{a\}^\{g\}\_\{i\}โs neighboring atom๐บjg\\mathsf\{a\}^\{g\}\_\{j\}in the ligand๐\\mathop\{\\mathcal\{D\}\}\\limits\(i\.e\.,๐\(๐บig\|๐\)\\mbox\{$\\mathop\{\\mathbf\{N\}\}\\limits$\}\(\\mathsf\{a\}^\{g\}\_\{i\}\|\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\)\) to๐บig\\mathsf\{a\}^\{g\}\_\{i\};zjโi,lz\_\{ji,l\}is the attention weights modeled in a similar way as in Equation[5](https://arxiv.org/html/2607.12349#Sx6.E5)\. The convolution in Equation[23](https://arxiv.org/html/2607.12349#Sx6.E23)combines information from ligand atom neighbors and pocket atom/residue neighbors, allowing each ligand atom to refine its representation with structural and chemical information from both the ligand and the pocket\.๐ฆjโi,lg\\mathbf\{m\}^\{g\}\_\{ji,l\}andMjโi,lgM^\{g\}\_\{ji,l\}are updated as follows:
\(๐ฆjโi,lg,Mjโi,lg\)=GVPโ\(๐ฆ^jโi,lg,M^jโi,lg\),where\(\\mathbf\{m\}^\{g\}\_\{ji,l\},M^\{g\}\_\{ji,l\}\)=\\text\{GVP\}\(\{\{\\hat\{\\mathbf\{m\}\}^\{g\}\_\{ji,l\}\}\},\{\{\\hat\{M\}^\{g\}\_\{ji,l\}\}\}\),\\text\{ where \}\(25\)๐ฆ^jโi,lg\\displaystyle\{\\hat\{\\mathbf\{m\}\}^\{g\}\_\{ji,l\}\}=\[๐ฆjโi,lโ1g,dโ\(๐บjg,๐บig\),๐jโi,lโ1\],\\displaystyle=\[\{\\mathbf\{m\}\}^\{g\}\_\{ji,l\-1\},\{d\(\\mathsf\{a\}^\{g\}\_\{j\},\\mathsf\{a\}^\{g\}\_\{i\}\)\},\_\{ji,l\-1\}\],\(26\)M^jโi,lg\\displaystyle\\quad\{\\hat\{M\}^\{g\}\_\{ji,l\}\}=\[Mjโi,lโ1g,๐ฑjgโ๐ฑig\],\\displaystyle=\[M^\{g\}\_\{ji,l\-1\},\{\\mathbf\{x\}^\{g\}\_\{j\}\-\\mathbf\{x\}^\{g\}\_\{i\}\}\],where๐jโi,lโ1\\mbox\{$\\mathop\{\\mathbf\{e\}\}\\limits$\}\_\{ji,l\-1\}is the embedding of the bond type between๐บig\\mathsf\{a\}^\{g\}\_\{i\}and๐บjg\\mathsf\{a\}^\{g\}\_\{j\}\(detailed in Equation[27](https://arxiv.org/html/2607.12349#Sx6.E27)\), anddโ\(๐บjg,๐บig\)d\(\\mathsf\{a\}^\{g\}\_\{j\},\\mathsf\{a\}^\{g\}\_\{i\}\)is the distance between๐บig\\mathsf\{a\}^\{g\}\_\{i\}and๐บjg\\mathsf\{a\}^\{g\}\_\{j\}, and\(๐ฑjgโ๐ฑig\)\(\\mathbf\{x\}^\{g\}\_\{j\}\-\\mathbf\{x\}^\{g\}\_\{i\}\)represents the displacement vector from๐บig\\mathsf\{a\}^\{g\}\_\{i\}to๐บjg\\mathsf\{a\}^\{g\}\_\{j\}\.๐๐ผ๐ซ๐ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitsutilizes bond type embeddings๐jโi,l\\mbox\{$\\mathop\{\\mathbf\{e\}\}\\limits$\}\_\{ji,l\}to facilitate its understanding of relations among atoms\. Particularly,๐๐ผ๐ซ๐ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitsgenerates these bond type embeddings as follows:
๐jโi,l=\{MLPโ\(\[๐ฌi,lg\+๐ฌj,lg,absโ\(๐ฌi,lgโ๐ฌj,lg\),dโ\(๐บjg,๐บig\)\]\),ifโl=0,MLPโ\(\[๐ฌi,lg\+๐ฌj,lg,absโ\(๐ฌi,lgโ๐ฌj,lg\),โโi,lgโ2\+โโj,lgโ2,absโ\(โโi,lgโ2โโโj,lgโ2\)\]\),otherwise\.\\mbox\{$\\mathop\{\\mathbf\{e\}\}\\limits$\}\_\{ji,l\}=\\begin\{cases\}\\text\{MLP\}\(\[\\mathbf\{s\}^\{g\}\_\{i,l\}\+\\mathbf\{s\}^\{g\}\_\{j,l\},\\text\{abs\}\(\\mathbf\{s\}^\{g\}\_\{i,l\}\-\\mathbf\{s\}^\{g\}\_\{j,l\}\),d\(\\mathsf\{a\}^\{g\}\_\{j\},\\mathsf\{a\}^\{g\}\_\{i\}\)\]\),&\\text\{if\}\\;l=0,\\\\ \\text\{MLP\}\(\[\\mathbf\{s\}^\{g\}\_\{i,l\}\+\\mathbf\{s\}^\{g\}\_\{j,l\},\\text\{abs\}\(\\mathbf\{s\}^\{g\}\_\{i,l\}\-\\mathbf\{s\}^\{g\}\_\{j,l\}\),\\\|\\mathcal\{H\}^\{g\}\_\{i,l\}\\\|\_\{2\}\+\\\|\\mathcal\{H\}^\{g\}\_\{j,l\}\\\|\_\{2\},\\text\{abs\}\(\\\|\\mathcal\{H\}^\{g\}\_\{i,l\}\\\|\_\{2\}\-\\\|\\mathcal\{H\}^\{g\}\_\{j,l\}\\\|\_\{2\}\)\]\),&\\text\{ otherwise\}\.\\end\{cases\}\(27\)where abs\(โ
\\cdot\) represents the absolute difference\. Intuitively, the bond type embeddings are constructed to be agnostic to the order of atoms๐บig\\mathsf\{a\}^\{g\}\_\{i\}and๐บjg\\mathsf\{a\}^\{g\}\_\{j\}by using two invariant operations: the sum and the absolute difference operation\. The sum combines features from neighboring atoms๐บig\\mathsf\{a\}^\{g\}\_\{i\}and๐บjg\\mathsf\{a\}^\{g\}\_\{j\}, and the absolute difference estimates the distance between๐บig\\mathsf\{a\}^\{g\}\_\{i\}and๐บjg\\mathsf\{a\}^\{g\}\_\{j\}on latent space, offering a comprehensive representation of the bond\. AfterLLlayers,๐๐ผ๐ซ๐ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitsestimates the atom position and type for each ligand atom๐บig\\mathsf\{a\}^\{g\}\_\{i\}using the learned atom embeddings as follows:
๐ฑ~ig=๐ฑig\+โi,Lg,๐ฏ~ig=softmaxโ\(MLPโ\(๐ฌi,Lg\)\),\{\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{i\}=\\mathbf\{x\}^\{g\}\_\{i\}\+\\mathcal\{H\}^\{g\}\_\{i,L\}\},\\quad\\tilde\{\\mathbf\{v\}\}^\{g\}\_\{i\}=\\text\{softmax\}\(\\text\{MLP\}\(\\mathbf\{s\}^\{g\}\_\{i,L\}\)\),\(28\)where๐ฑ~ig\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{i\}and๐ฑig\\mathbf\{x\}^\{g\}\_\{i\}are the predicted noise\-free atom positions and input noisy atom positions of atom๐บig\\mathsf\{a\}^\{g\}\_\{i\}, respectively;โi,Lgโโ1ร3\\mathcal\{H\}^\{g\}\_\{i,L\}\\in\\mathbb\{R\}^\{1\\times 3\}is the predicted noise for๐บig\\mathsf\{a\}^\{g\}\_\{i\}, and๐ฌi,Lg\\mathbf\{s\}^\{g\}\_\{i,L\}is used for atom type prediction of๐บig\\mathsf\{a\}^\{g\}\_\{i\}\.๐ฏ~ig\\tilde\{\\mathbf\{v\}\}^\{g\}\_\{i\}represents the predicted clean categorical distribution of atom features\.
#### Pocket\-ligand Interaction Learning
This section describes how to learn pocket\-ligand interaction embeddings \(embeddings with superscriptcorrin Equation[24](https://arxiv.org/html/2607.12349#Sx6.E24)\) by leveraging pocket embeddings \(embeddings with superscriptp\) encoded by๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits\. For each ligand atom๐บig\\mathsf\{a\}^\{g\}\_\{i\}, at eachll\-th layer of the GNN,๐๐ผ๐ซ๐ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitslearns two invariant embeddings๐ฌi,lc\\mathbf\{s\}^\{c\}\_\{i,l\}and๐ฌi,la\\mathbf\{s\}^\{a\}\_\{i,l\}, and two equivariant embeddingโi,lc\\mathcal\{H\}^\{c\}\_\{i,l\}andโi,la\\mathcal\{H\}^\{a\}\_\{i,l\}, encoding how๐บig\\mathsf\{a\}^\{g\}\_\{i\}interacts with neighboring pocket atoms and residues\. These embeddings are learned in a recurrent manner as follows:
๐ฌi,lc=MLPโ\(\[๐ฌ~i,lc,๐ฌi,lโ1c\]\),โi,lc=VN\-MLPโ\(\[โ~i,lc,โi,lโ1c\]\),\\mathbf\{s\}^\{c\}\_\{i,l\}=\\text\{MLP\}\(\[\\tilde\{\\mathbf\{s\}\}^\{c\}\_\{i,l\},\\mathbf\{s\}^\{c\}\_\{i,l\-1\}\]\),~~\\mathcal\{H\}^\{c\}\_\{i,l\}=\\text\{VN\-MLP\}\(\[\\tilde\{\\mathcal\{H\}\}^\{c\}\_\{i,l\},\\mathcal\{H\}^\{c\}\_\{i,l\-1\}\]\),\(29\)๐ฌi,la=MLPโ\(\[๐ฌ~i,la,๐ฌi,lโ1a\]\),โi,la=VN\-MLPโ\(\[โ~i,la,โi,lโ1a\]\)\.\\mathbf\{s\}^\{a\}\_\{i,l\}=\\text\{MLP\}\(\[\\tilde\{\\mathbf\{s\}\}^\{a\}\_\{i,l\},\\mathbf\{s\}^\{a\}\_\{i,l\-1\}\]\),~~\\mathcal\{H\}^\{a\}\_\{i,l\}=\\text\{VN\-MLP\}\(\[\\tilde\{\\mathcal\{H\}\}^\{a\}\_\{i,l\},\\mathcal\{H\}^\{a\}\_\{i,l\-1\}\]\)\.\(30\)That is,๐๐ผ๐ซ๐ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitsupdates the atom\-level\(๐ฌic,โic\)\(\\mathbf\{s\}^\{c\}\_\{i\},\\mathcal\{H\}^\{c\}\_\{i\}\)and residue\-level\(๐ฌia,โia\)\(\\mathbf\{s\}^\{a\}\_\{i\},\\mathcal\{H\}^\{a\}\_\{i\}\)pocket\-ligand interaction embeddings recurrently across layers\. We model๐ฌ~i,lc\\tilde\{\\mathbf\{s\}\}^\{c\}\_\{i,l\}andโ~i,lc\\tilde\{\\mathcal\{H\}\}^\{c\}\_\{i,l\}as follows:
\(๐ฌ~i,lc,โ~i,lc\)=GVPโ\(๐i,lc,โi,lc\),where\(\\tilde\{\\mathbf\{s\}\}^\{c\}\_\{i,l\},\\tilde\{\\mathcal\{H\}\}^\{c\}\_\{i,l\}\)=\\text\{GVP\}\(\\mathbf\{b\}\_\{i,l\}^\{c\},\\mathcal\{I\}\_\{i,l\}^\{c\}\),\\text\{ where \}\(31\)๐i,lc\\displaystyle\\mathbf\{b\}\_\{i,l\}^\{c\}=โjโ๐\(๐บig\|๐ซ\)nwiโj,lcโMLPโ\(\[dโ\(๐บig,๐บjp\),๐ฌi,lg,๐ฌjp\]\),\\displaystyle=\\sum\_\{j\\in\{\\mbox\{$\\mathop\{\\mathbf\{N\}\}\\limits$\}\(\\mathsf\{a\}^\{g\}\_\{i\}\|\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\)\}\}^\{n\}w\_\{ij,l\}^\{c\}\\text\{MLP\}\(\[d\(\\mathsf\{a\}^\{g\}\_\{i\},\\mathsf\{a\}^\{p\}\_\{j\}\),\\mathbf\{s\}^\{g\}\_\{i,l\},\\mathbf\{s\}^\{p\}\_\{j\}\]\),\(32\)โi,lc\\displaystyle\\mathcal\{I\}\_\{i,l\}^\{c\}=โ๐บjpโ๐\(๐บig\|๐ซ\)nwiโj,lcโVN\-MLPโ\(\[๐ฑjpโ๐ฑig,โi,lg,โjp\]\),\\displaystyle=\\sum\_\{\\mathsf\{a\}^\{p\}\_\{j\}\\in\{\\mbox\{$\\mathop\{\\mathbf\{N\}\}\\limits$\}\(\\mathsf\{a\}^\{g\}\_\{i\}\|\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\)\}\}^\{n\}w\_\{ij,l\}^\{c\}\\text\{VN\-MLP\}\(\[\\mathbf\{x\}^\{p\}\_\{j\}\-\\mathbf\{x\}^\{g\}\_\{i\},\\mathcal\{H\}^\{g\}\_\{i,l\},\\mathcal\{H\}^\{p\}\_\{j\}\]\),where๐ซ\\mathop\{\\mathcal\{P\}\}\\limitsis the target pocket of ligand๐\\mathop\{\\mathcal\{D\}\}\\limits, and๐\(๐บig\|๐ซ\)\{\\mbox\{$\\mathop\{\\mathbf\{N\}\}\\limits$\}\(\\mathsf\{a\}^\{g\}\_\{i\}\|\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\)\}denotes thenn\-nearest pocket atoms surrounding the ligand atom๐บig\\mathsf\{a\}^\{g\}\_\{i\};๐ฌjp\\mathbf\{s\}^\{p\}\_\{j\}andโjp\\mathcal\{H\}^\{p\}\_\{j\}are pretrained pocket atom embeddings \(Equation[3](https://arxiv.org/html/2607.12349#Sx6.E3)\);๐ฌi,lโ1g\\mathbf\{s\}^\{g\}\_\{i,l\-1\}andโi,lโ1g\\mathcal\{H\}^\{g\}\_\{i,l\-1\}are ligand atom embeddings as presented in Equation[23](https://arxiv.org/html/2607.12349#Sx6.E23); anddโ\(๐บig,๐บjp\)d\(\\mathsf\{a\}^\{g\}\_\{i\},\\mathsf\{a\}^\{p\}\_\{j\}\)is the Euclidean distance between๐บig\\mathsf\{a\}^\{g\}\_\{i\}and๐บjp\\mathsf\{a\}^\{p\}\_\{j\}\. The attention weights are determined based on the ligand embeddings \(Equation[23](https://arxiv.org/html/2607.12349#Sx6.E23)\) and pocket embeddings provided by๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limitsas follows:
wiโj,lc\\displaystyle w\_\{ij,l\}^\{c\}=softmaxโ\(MLPโ\(\[dโ\(๐บig,๐บjp\),๐ฏiโj,lc,๐ฌi,lg,๐ฌjp\]\)\),\\displaystyle=\\text\{softmax\}\(\\text\{MLP\}\(\[d\(\\mathsf\{a\}^\{g\}\_\{i\},\\mathsf\{a\}^\{p\}\_\{j\}\),\\mathbf\{v\}^\{c\}\_\{ij,l\},\\mathbf\{s\}^\{g\}\_\{i,l\},\\mathbf\{s\}^\{p\}\_\{j\}\]\)\),\(33\)๐ฏiโj,lc\\displaystyle\\mathbf\{v\}\_\{ij,l\}^\{c\}=โVN\-MLPโ\(\[๐ฑjpโ๐ฑig,โi,lg,โjp\]\)โ2,\\displaystyle=\\\|\\text\{VN\-MLP\}\(\[\\mathbf\{x\}^\{p\}\_\{j\}\-\\mathbf\{x\}^\{g\}\_\{i\},\\mathcal\{H\}^\{g\}\_\{i,l\},\\mathcal\{H\}^\{p\}\_\{j\}\]\)\\\|\_\{2\},where๐ฏiโj,l\\mathbf\{v\}\_\{ij,l\}encodes the spatial interactions between๐บig\\mathsf\{a\}^\{g\}\_\{i\}and๐บjp\\mathsf\{a\}^\{p\}\_\{j\}\. While\(๐ฌ~i,lc,โ~i,lc\)\(\\tilde\{\\mathbf\{s\}\}^\{c\}\_\{i,l\},\\tilde\{\\mathcal\{H\}\}^\{c\}\_\{i,l\}\)accounts for atomic pocket\-ligand interactions, we also model interactions between ligand atoms and pocket residues\.\(๐ฌ~i,la,โ~i,la\)\(\\tilde\{\\mathbf\{s\}\}^\{a\}\_\{i,l\},\\tilde\{\\mathcal\{H\}\}^\{a\}\_\{i,l\}\)are obtained in a similar manner by aggregating interactions between๐บig\\mathsf\{a\}^\{g\}\_\{i\}andmm\-nearest pocket residue๐jp\\mathsf\{r\}^\{p\}\_\{j\}\. We summarize the ligand generation procedure of๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsincorporating๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limitsand๐๐ผ๐ซ๐ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitsin Algorithm[2](https://arxiv.org/html/2607.12349#alg2)\.
Algorithm 2๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsfor ligand generationRequired Input: Ligand๐\\mathop\{\\mathcal\{D\}\}\\limits, pocket๐ซ\\mathop\{\\mathcal\{P\}\}\\limits
1:
n=sampleNumAtomsโ\(๐,๐ซ\)n=\\text\{\{sampleNumAtoms\}\}\(\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\},\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\)โณ\\trianglerightsample the number of ligand atoms from pocket size
2:
\{๐ฑTg\}nโผ๐ฉโ\(0,๐\)\\\{\\mathbf\{x\}^\{g\}\_\{T\}\\\}^\{n\}\\sim\\mathcal\{N\}\(0,\\mathbf\{I\}\)โณ\\trianglerightinitialize positions of n ligand atoms
3:
\{๐ฏTg\}nโผ๐โ\(K,1K\)\\\{\\mathbf\{v\}^\{g\}\_\{T\}\\\}^\{n\}\\sim\\mathcal\{C\}\(K,\\frac\{1\}\{K\}\)โณ\\trianglerightinitialize types of n ligand atoms
4:
๐ฌp,โp=๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\(๐ฑp,๐ฏp,๐ฑr,๐ฏr\)\\mathbf\{s\}^\{p\},\\mathcal\{H\}^\{p\}=\\mbox\{$\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits$\}\(\\mathbf\{x\}^\{p\},\\mathbf\{v\}^\{p\},\\mathbf\{x\}^\{r\},\\mathbf\{v\}^\{r\}\)โณ\\trianglerightencode pocket into embeddings using๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits
5:for
t=Tt=Tto
11do
6:
\(๐ฑ~0,tg,๐ฑ~0,tg\)=๐๐ผ๐ซ๐ฆ\(๐ฑtg,๐ฏtg,๐ฌp,โp\)\(\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{0,t\},\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{0,t\}\)=\\mbox\{$\\mathop\{\\mathsf\{pcLG\}\}\\limits$\}\(\\mathbf\{x\}^\{g\}\_\{t\},\\mathbf\{v\}^\{g\}\_\{t\},\\mathbf\{s\}^\{p\},\\mathcal\{H\}^\{p\}\)โณ\\trianglerightpredict noise\-free ligand using๐๐ผ๐ซ๐ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limits
7:
๐ฑtโ1g=qโ\(๐ฑtโ1g\|๐ฑtg,๐ฑ~0,tg\)\\mathbf\{x\}^\{g\}\_\{t\-1\}=q\(\\mathbf\{x\}^\{g\}\_\{t\-1\}\|\\mathbf\{x\}^\{g\}\_\{t\},\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{0,t\}\)โณ\\trianglerightsample๐ฑtโ1g\\mathbf\{x\}^\{g\}\_\{t\-1\}using Gaussian posterior \(Equation[16](https://arxiv.org/html/2607.12349#Sx6.E16)\)
8:
๐ฏtโ1g=qโ\(๐ฏtโ1g\|๐ฏtg,๐ฑ~0,tg\)\\mathbf\{v\}^\{g\}\_\{t\-1\}=q\(\\mathbf\{v\}^\{g\}\_\{t\-1\}\|\\mathbf\{v\}^\{g\}\_\{t\},\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{0,t\}\)โณ\\trianglerightsample๐ฏtโ1g\\mathbf\{v\}^\{g\}\_\{t\-1\}using categorical posterior \(Equation[17](https://arxiv.org/html/2607.12349#Sx6.E17)\)
9:endfor
10:
๐gโeโn=\(๐ฑ0g,๐ฏ0g\)\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\_\{gen\}=\(\\mathbf\{x\}^\{g\}\_\{0\},\\mathbf\{v\}^\{g\}\_\{0\}\)
11:return
๐gโeโn\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\_\{gen\}โณ\\trianglerightreturn the generated ligand
### ๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitswith Generation\-time Property\-aware Optimization \(๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\)
๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsis well tailored for pocket\-specific ligand generation\. In this section, we develop a new method that further improves the ligands generated by๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsto exhibit additional desired properties \(e\.g\., ADMET\)\. To this end, we extend๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsby introducing an iterative, generation\-time property\-aware optimization procedure that refines the generated ligands with respect to these desired properties\. In essence, this optimization process โalignsโ the model with desired properties beyond binding affinity or drug\-likeness\. We formulate this โalignmentโ process as a generation\-time optimization problem, in which the objective is to steer the generated ligands toward exhibiting more favorable properties, such as higher BBBP or lower carcinogenicity\. The proposed optimization scheme, named๐๐บ๐ฎ๐ฏ๐ณ\\mathop\{\\mathsf\{paOPT\}\}\\limits, operates entirely during the modelโs inference process\. That is, all model parameters of๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsare kept fixed at this stage\. Without requiring any additional model training, the approach is both computationally efficient and readily adaptable to different optimization objectives while preserving๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsโs original generative performance on general molecular metrics\. We refer to the resulting framework that combines๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitswith๐๐บ๐ฎ๐ฏ๐ณ\\mathop\{\\mathsf\{paOPT\}\}\\limitsas๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\.
#### Alignment Problem Formulation
In๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limits, the stochasticity of ligand generation arises from noise trajectories of atom positions and atom types \(๐ณt๐ฑ\\mathbf\{z\}^\{\\mathbf\{x\}\}\_\{t\}in Equation[16](https://arxiv.org/html/2607.12349#Sx6.E16)and๐ณt๐ฏ\\mathbf\{z\}^\{\\mathbf\{v\}\}\_\{t\}in Equation[17](https://arxiv.org/html/2607.12349#Sx6.E17)\)\. Noise trajectories govern the denoising dynamics throughout the diffusion process, determining how ligand structures evolve over time and ultimately shaping the final generated ligand\.
Therefore, with the diffusion model parameters fixed, the entire generation process can be viewed as a deterministic mapping that takes the noise trajectories as input and outputs the generated ligand\. These noise trajectories constitute the optimization variables in๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\. After generation, the ligand is evaluated by an evaluator functionRโ\(โ
\)R\(\\cdot\), which defines the objective function for the optimization problem\. We then employ gradient\-based optimization to iteratively update the noise trajectories, using the gradients of the objective function to steer the generation process towards ligands with more favorable properties\.
##### Noise as Optimization Variables
The generation process of๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limits, described in Algorithm[2](https://arxiv.org/html/2607.12349#alg2), consists ofTTdenoising steps\. It begins with atom positions initialized as pure Gaussian noise,๐ฑi,Tgโผ๐ฉโ\(๐,๐\)\\mathbf\{x\}^\{g\}\_\{i,T\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)\(here, we define๐ณi,T๐ก=๐ฑi,Tg\\mathbf\{z\}^\{\\mathtt\{x\}\}\_\{i,T\}=\\mathbf\{x\}^\{g\}\_\{i,T\}\), and atom types sampled uniformly at random,๐ฏi,TgโผCโ\(K,1K\)\.\\mathbf\{v\}^\{g\}\_\{i,T\}\\sim C\(K,\\frac\{1\}\{K\}\)\.The uniform categorical sampling is implemented via the Gumbel\-max trick\. Specifically, for each atomiiand typekk, in๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits, we draw๐ณi,T,k๐โผGumbelโ\(0,1\)\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{i,T,k\}\\sim\\mathrm\{Gumbel\}\(0,1\)fork=1,โฆ,Kk=1,\\dots,K, and set๐ฏi,Tg=argโกmaxkโก\{๐ณi,T,k๐\}\\mathbf\{v\}^\{g\}\_\{i,T\}=\\arg\\max\_\{k\}\\\{\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{i,T,k\}\\\}\. At each subsequent steptt, the atom positions๐ฑi,tg\\mathbf\{x\}^\{g\}\_\{i,t\}and types๐ฏi,tg\\mathbf\{v\}^\{g\}\_\{i,t\}are updated according to Equations[16](https://arxiv.org/html/2607.12349#Sx6.E16)and[17](https://arxiv.org/html/2607.12349#Sx6.E17), where Gaussian noise๐ณi,tโ1๐กโผ๐ฉโ\(๐,๐\)\\mathbf\{z\}^\{\\mathtt\{x\}\}\_\{i,t\-1\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)and Gumbel noises๐ณi,tโ1,k๐โผGumbel\(0,1\)\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{i,t\-1,k\}\\sim\\text\{Gumbel\(0,1\)\}fork=1,โฆ,Kk=1,\\dots,Kare injected into the denoising updates to produce the predicted๐ฑi,tโ1g\\mathbf\{x\}^\{g\}\_\{i,t\-1\}and๐ฏi,tโ1g\\mathbf\{v\}^\{g\}\_\{i,t\-1\}\. For compactness, we write\[๐ณi,t,1๐,โฆ,๐ณi,t,k๐\]\[\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{i,t,1\},\\dots,\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{i,t,k\}\]as๐ณi,t๐\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{i,t\}\. We collect these noises into two noise trajectories: Gaussian noise trajectory\{๐ณi,T๐ก,๐ณi,Tโ1๐ก,โฆ,๐ณi,0๐ก\}\\\{\\mathbf\{z\}^\{\\mathtt\{x\}\}\_\{i,T\},\\mathbf\{z\}^\{\\mathtt\{x\}\}\_\{i,T\-1\},\\dots,\\mathbf\{z\}^\{\\mathtt\{x\}\}\_\{i,0\}\\\}and Gumbel noise trajectory\{๐ณi,T๐,๐ณi,T๐,โฆ,๐ณi,0๐\}\\\{\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{i,T\},\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{i,T\},\\dots,\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{i,0\}\\\}\. By treating the noise trajectories as optimization variables, we aim to iteratively refine them to guide the generative process towards chemical space with improved target properties\.
##### Objective Function
LetD0D\_\{0\}represent the generated ligand at the final timestep\. To evaluate its quality with respect to a desired ADMET property, we employ a property\-specific evaluator function\. We denote the evaluator asRโ\(โ
\)R\(\\cdot\), which serves as the objective function in our formulation\. The choice of evaluator is flexible and can be tailored to specific target properties\. For example, if the goal is to improve the blood\-brain barrier permeability of the generated ligandD0D\_\{0\}, thenRโ\(D0\)R\(D\_\{0\}\)can be defined as the predicted probability \(given by an external property classifier\) that ligandD0D\_\{0\}is blood\-brain barrier permeable\.
##### Noise Optimization
For notational simplicity, we denote the entire Gaussian noise trajectory\{๐ณi,T๐ก,๐ณi,Tโ1๐ก,โฆ,๐ณi,0๐ก\}\\\{\\mathbf\{z\}^\{\\mathtt\{x\}\}\_\{i,T\},\\mathbf\{z\}^\{\\mathtt\{x\}\}\_\{i,T\-1\},\\dots,\\mathbf\{z\}^\{\\mathtt\{x\}\}\_\{i,0\}\\\}asฯ๐ฑ\\tau^\{\\mathbf\{x\}\}and the entire Gumbel noise trajectory\{๐ณi,T๐,๐ณi,Tโ1๐,โฆ,๐ณi,0๐\}\\\{\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{i,T\},\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{i,T\-1\},\\dots,\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{i,0\}\\\}asฯ๐ฏ\\tau^\{\\mathbf\{v\}\}\. With๐๐ผ๐ซ๐ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitsfixed,ฯ๐ฑ\\tau^\{\\mathbf\{x\}\}andฯ๐ฏ\\tau^\{\\mathbf\{v\}\}fully determine the generated ligand๐0\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\_\{0\}\. Therefore, we write๐0=Mโ\(ฯ๐ฑ,ฯ๐ฏ\)\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\_\{0\}=M\(\\tau^\{\\mathbf\{x\}\},\\tau^\{\\mathbf\{v\}\}\), whereMMdenotes the deterministic mapping from noise trajectories to the generated ligand\. Our goal is to adjust the noise trajectoriesฯ๐ฑ\\tau^\{\\mathbf\{x\}\}andฯ๐ฏ\\tau^\{\\mathbf\{v\}\}so that the generated ligand has a more favorable property scoreRR\. The generation\-time noise optimization problem is thus formulated as:
maxฯ๐ฑ,ฯ๐ฏโกRโ\(Mโ\(ฯ๐ฑ,ฯ๐ฏ\)\)\\max\_\{\\tau^\{\\mathbf\{x\}\},\\tau^\{\\mathbf\{v\}\}\}R\(M\(\\tau^\{\\mathbf\{x\}\},\\tau^\{\\mathbf\{v\}\}\)\)\(34\)
#### ๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsOptimization
We adopt a straightforward gradient\-based method, as presented in Algorithm[3](https://arxiv.org/html/2607.12349#alg3), to solve the noise optimization problem \(Equation[34](https://arxiv.org/html/2607.12349#Sx6.E34)\)\. We use the noise trajectories produced by๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsas a starting point to be optimized\. The initial noise trajectories are iteratively refined through gradient\-based updates\. Specifically, at each iteration, we evaluate the objectiveRโ\(Mโ\(ฯ๐ฑ,ฯ๐ฏ\)\)R\(M\(\\tau^\{\\mathbf\{x\}\},\\tau^\{\\mathbf\{v\}\}\)\), calculate the gradient ofRRwith respect toฯ๐ฑ,ฯ๐ฏ\\tau^\{\\mathbf\{x\}\},\\tau^\{\\mathbf\{v\}\}, and update the noise accordingly via gradient ascent\. This approach is applicable when the gradient of the objective functionRRis available\. However, in practice, such gradients are often inaccessible, since many ADMET evaluators operate as black\-box predictors that provide only scalar property scores\. To overcome this limitation, we employ zero*th*\-order gradient estimation, enabling gradient\-based updates without explicit analytic gradients\. Algorithm[5](https://arxiv.org/html/2607.12349#alg5)presents the๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsalgorithm\.
Algorithm 3๐ญ๐๐๐๐พ๐ฎ๐๐\\mathop\{\\mathsf\{NoiseOpt\}\}\\limitsfor noise trajectory optimizationRequired Input: Initial noise trajectoriesฯ0๐ฑ,ฯ0๐ฏ\\tau^\{\\mathbf\{x\}\}\_\{0\},\\tau^\{\\mathbf\{v\}\}\_\{0\}, iteration numberNN, step sizeฮฑ\\alpha
1:
rbestโRโ\(Mโ\(ฯ0๐ฑ,ฯ0๐ฏ\)\),nbestโ0r\_\{\\mathrm\{best\}\}\\leftarrow R\(M\(\\tau^\{\\mathbf\{x\}\}\_\{0\},\\tau^\{\\mathbf\{v\}\}\_\{0\}\)\),n\_\{\\mathrm\{best\}\}\\leftarrow 0โณ\\trianglerightevaluate objective at initialization
2:for
n=0n=0to
Nโ1N\-1do
3:
rโRโ\(Mโ\(ฯn๐ฑ,ฯn๐ฏ\)\)r\\leftarrow R\(M\(\\tau^\{\\mathbf\{x\}\}\_\{n\},\\tau^\{\\mathbf\{v\}\}\_\{n\}\)\)โณ\\trianglerightevaluate current objective
4:if
r\>rbestr\>r\_\{\\mathrm\{best\}\}then
5:
rbestโr,nbestโnr\_\{\\mathrm\{best\}\}\\leftarrow r,n\_\{\\mathrm\{best\}\}\\leftarrow nโณ\\trianglerightupdate best step
6:endif
7:
โ^โRโ\(Mโ\(ฯn๐ฑ,ฯn๐ฏ\)\)=๐น๐ฎ๐ฆ๐๐บ๐ฝ\(ฯn๐ฑ,ฯn๐ฏ\)\\hat\{\\nabla\}R\(M\(\\tau^\{\\mathbf\{x\}\}\_\{n\},\\tau^\{\\mathbf\{v\}\}\_\{n\}\)\)=\\mbox\{$\\mathop\{\\mathsf\{ZOGrad\}\}\\limits$\}\(\\tau^\{\\mathbf\{x\}\}\_\{n\},\\tau^\{\\mathbf\{v\}\}\_\{n\}\)โณ\\trianglerightestimate zeroth\-order gradients via Algorithm[4](https://arxiv.org/html/2607.12349#alg4)
8:
ฯn\+1๐ฑ=ฯn๐ฑ\+ฮฑโโ^ฯ๐ฑโRโ\(Mโ\(ฯn๐ฑ,ฯn๐ฏ\)\)\\tau^\{\\mathbf\{x\}\}\_\{n\+1\}=\\tau^\{\\mathbf\{x\}\}\_\{n\}\+\\alpha\\hat\{\\nabla\}\_\{\\tau^\{\\mathbf\{x\}\}\}R\(M\(\\tau^\{\\mathbf\{x\}\}\_\{n\},\\tau^\{\\mathbf\{v\}\}\_\{n\}\)\)โณ\\trianglerightupdate Gaussian noise trajectory
9:
ฯn\+1๐ฏ=ฯn๐ฏ\+ฮฑโโ^ฯ๐ฏโRโ\(Mโ\(ฯn๐ฑ,ฯn๐ฏ\)\)\\tau^\{\\mathbf\{v\}\}\_\{n\+1\}=\\tau^\{\\mathbf\{v\}\}\_\{n\}\+\\alpha\\hat\{\\nabla\}\_\{\\tau^\{\\mathbf\{v\}\}\}R\(M\(\\tau^\{\\mathbf\{x\}\}\_\{n\},\\tau^\{\\mathbf\{v\}\}\_\{n\}\)\)โณ\\trianglerightupdate Gumbel noise trajectory
10:endfor
11:
๐bestโMโ\(ฯnbest๐ฑ,ฯnbest๐ฏ\)\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\_\{\\mathrm\{best\}\}\\leftarrow M\(\\tau^\{\\mathbf\{x\}\}\_\{n\_\{\\mathrm\{best\}\}\},\\tau^\{\\mathbf\{v\}\}\_\{n\_\{\\mathrm\{best\}\}\}\)
12:return
๐best\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\_\{\\mathrm\{best\}\}โณ\\trianglerightreturn the best ligand found among all iterations
##### Zero*th*\-order Gradient Estimation
Algorithm[4](https://arxiv.org/html/2607.12349#alg4)outlines the zero*th*\-order gradient estimation procedure used in๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits\. The key idea is to estimate the gradient of the objective function by observing how small random perturbations to the input variables change the output score\. At each iteration, we sample random perturbation directions for both the Gaussian and Gumbel noise trajectories,u๐ฑโผ๐ฉโ\(0,๐๐ฑ\)u^\{\\mathbf\{x\}\}\\sim\\mathcal\{N\}\(0,\\mathbf\{I\}^\{\\mathbf\{x\}\}\)andu๐ฏโผ๐ฉโ\(0,๐๐ฏ\)u^\{\\mathbf\{v\}\}\\sim\\mathcal\{N\}\(0,\\mathbf\{I\}^\{\\mathbf\{v\}\}\), and generate perturbed versions of the trajectories:\(ฯ๐ฑ\+ฮผโu๐ฑ,ฯ๐ฏ\+ฮผโu๐ฏ\)\(\\tau^\{\\mathbf\{x\}\}\+\\mu u^\{\\mathbf\{x\}\},\\tau^\{\\mathbf\{v\}\}\+\\mu u^\{\\mathbf\{v\}\}\)and\(ฯ๐ฑโฮผโu๐ฑ,ฯ๐ฏโฮผโu๐ฏ\)\(\\tau^\{\\mathbf\{x\}\}\-\\mu u^\{\\mathbf\{x\}\},\\tau^\{\\mathbf\{v\}\}\-\\mu u^\{\\mathbf\{v\}\}\), whereฮผ\\muis the perturbation magnitude\. The objective function is then evaluated at these perturbed points, and the gradient is estimated from the scaled difference between the corresponding objective values, providing a stochastic estimation of the true gradient direction\. By averaging the estimated directional derivatives over multiple perturbations, we obtain unbiased gradient estimates with respect to bothฯ๐ฑ\\tau^\{\\mathbf\{x\}\}andฯ๐ฏ\\tau^\{\\mathbf\{v\}\}\. This enables effective gradient\-based optimization even when the evaluator is a black\-box predictor that provides only function values without analytic gradients\.
Algorithm 4๐น๐ฎ๐ฆ๐๐บ๐ฝ\\mathop\{\\mathsf\{ZOGrad\}\}\\limitsfor zeroth\-order gradient estimationRequired Input: Current trajectories\(ฯ๐ฑ,ฯ๐ฏ\)\(\\tau^\{\\mathbf\{x\}\},\\tau^\{\\mathbf\{v\}\}\), smoothingฮผ\>0\\mu\>0, number of random perturbationsHH
1:Initialize accumulators:
g๐ฑโ๐g^\{\\mathbf\{x\}\}\\\!\\leftarrow\\mathbf\{0\},
g๐ฏโ๐g^\{\\mathbf\{v\}\}\\\!\\leftarrow\\mathbf\{0\}
2:for
h=1h=1to
HHdo
3:
uh๐ฑโผ๐ฉโ\(0,๐๐ฑ\)u^\{\\mathbf\{x\}\}\_\{h\}\\sim\\mathcal\{N\}\(0,\\mathbf\{I^\{\\mathbf\{x\}\}\}\),
uh๐ฏโผ๐ฉโ\(0,๐๐ฏ\)u^\{\\mathbf\{v\}\}\_\{h\}\\sim\\mathcal\{N\}\(0,\\mathbf\{I^\{\\mathbf\{v\}\}\}\)โณ\\trianglerightsample random perturbations
4:
rh\+โRโ\(Mโ\(ฯ๐ฑ\+ฮผโuh๐ฑ,ฯ๐ฏ\+ฮผโuh๐ฏ\)\)r^\{\+\}\_\{h\}\\leftarrow R\\\!\\left\(M\\\!\\big\(\\tau^\{\\mathbf\{x\}\}\+\\mu u^\{\\mathbf\{x\}\}\_\{h\},\\;\\tau^\{\\mathbf\{v\}\}\+\\mu u^\{\\mathbf\{v\}\}\_\{h\}\\big\)\\right\)โณ\\trianglerightpositive perturbations
5:
rhโโRโ\(Mโ\(ฯ๐ฑโฮผโuh๐ฑ,ฯ๐ฏโฮผโuh๐ฏ\)\)r^\{\-\}\_\{h\}\\leftarrow R\\\!\\left\(M\\\!\\big\(\\tau^\{\\mathbf\{x\}\}\-\\mu u^\{\\mathbf\{x\}\}\_\{h\},\\;\\tau^\{\\mathbf\{v\}\}\-\\mu u^\{\\mathbf\{v\}\}\_\{h\}\\big\)\\right\)โณ\\trianglerightnegative perturbations
6:
ฮโrhโrh\+โrhโ2โฮผ\\Delta r\_\{h\}\\leftarrow\\dfrac\{r^\{\+\}\_\{h\}\-r^\{\-\}\_\{h\}\}\{2\\mu\}โณ\\trianglerightcalculate directional difference
7:
g๐ฑโg๐ฑ\+ฮโrhโuh๐ฑ,g๐ฏโg๐ฏ\+ฮโrhโuh๐ฏg^\{\\mathbf\{x\}\}\\leftarrow g^\{\\mathbf\{x\}\}\+\\Delta r\_\{h\}\\,u^\{\\mathbf\{x\}\}\_\{h\},\\qquad g^\{\\mathbf\{v\}\}\\leftarrow g^\{\\mathbf\{v\}\}\+\\Delta r\_\{h\}\\,u^\{\\mathbf\{v\}\}\_\{h\}โณ\\trianglerightaccumulate estimators
8:endfor
9:
โ^ฯ๐ฑโRโ\(Mโ\(ฯ๐ฑ,ฯ๐ฏ\)\)โ1Hโg๐ฑ\\widehat\{\\nabla\}\_\{\\tau^\{\\mathbf\{x\}\}\}R\(M\(\\tau^\{\\mathbf\{x\}\},\\tau^\{\\mathbf\{v\}\}\)\)\\leftarrow\\dfrac\{1\}\{H\}g^\{\\mathbf\{x\}\},
โ^ฯ๐ฏโRโ\(Mโ\(ฯ๐ฑ,ฯ๐ฏ\)\)โ1Hโg๐ฏ\\widehat\{\\nabla\}\_\{\\tau^\{\\mathbf\{v\}\}\}R\(M\(\\tau^\{\\mathbf\{x\}\},\\tau^\{\\mathbf\{v\}\}\)\)\\leftarrow\\dfrac\{1\}\{H\}g^\{\\mathbf\{v\}\}
10:Return
โ^ฯ๐ฑโRโ\(Mโ\(ฯ๐ฑ,ฯ๐ฏ\)\)\\widehat\{\\nabla\}\_\{\\tau^\{\\mathbf\{x\}\}\}R\(M\(\\tau^\{\\mathbf\{x\}\},\\tau^\{\\mathbf\{v\}\}\)\),
โ^ฯ๐ฏโRโ\(Mโ\(ฯ๐ฑ,ฯ๐ฏ\)\)\\widehat\{\\nabla\}\_\{\\tau^\{\\mathbf\{v\}\}\}R\(M\(\\tau^\{\\mathbf\{x\}\},\\tau^\{\\mathbf\{v\}\}\)\)
Algorithm 5๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsfor ligand generation and optimizationRequired Input:๐\\mathop\{\\mathcal\{D\}\}\\limits,๐ซ\\mathop\{\\mathcal\{P\}\}\\limits, iteration numberNN
1:
n=sampleNumAtomsโ\(๐,๐ซ\)n=\\text\{\{sampleNumAtoms\}\}\(\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\},\\mbox\{$\\mathop\{\\mathcal\{P\}\}\\limits$\}\)โณ\\trianglerightsample the number of ligand atoms from pocket size
2:
\{๐ณi,T๐ก\}i=1nโผ๐ฉโ\(0,๐\),\{๐ฑTg\}i=1n=\{๐ณT๐ก\}i=1n\\\{\\mathbf\{z\}^\{\\mathtt\{x\}\}\_\{i,T\}\\\}\_\{i=1\}^\{n\}\\sim\\mathcal\{N\}\(0,\\mathbf\{I\}\),\\quad\\\{\\mathbf\{x\}^\{g\}\_\{T\}\\\}\_\{i=1\}^\{n\}=\\\{\\mathbf\{z\}^\{\\mathtt\{x\}\}\_\{T\}\\\}\_\{i=1\}^\{n\}โณ\\trianglerightinitialize positions ofnnligand atoms
3:
\{๐ณi,T๐\}i=1nโผGumbelโ\(0,1\)K,\{๐ฏTg\}i=1n=\{onehotโ\(argโกmaxkโ\[K\]โก๐ณi,T,k๐\)\}i=1n\\\{\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{i,T\}\\\}\_\{i=1\}^\{n\}\\sim\\mathrm\{Gumbel\}\(0,1\)^\{K\},\\quad\\\{\\mathbf\{v\}^\{g\}\_\{T\}\\\}\_\{i=1\}^\{n\}=\\left\\\{\\mathrm\{onehot\}\\\!\\left\(\\arg\\max\_\{k\\in\[K\]\}\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{i,T,k\}\\right\)\\right\\\}\_\{i=1\}^\{n\}โณ\\trianglerightinitialize types ofnnligand atoms
4:
ฯ๐ฑโโ
,ฯ๐ฏโโ
\\tau^\{\\mathbf\{x\}\}\\leftarrow\\emptyset,\\;\\tau^\{\\mathbf\{v\}\}\\leftarrow\\emptysetโณ\\trianglerightinitialize noise trajectories
5:
๐ฌp,โp=๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\(๐ฑp,๐ฏp,๐ฑr,๐ฏr\)\\mathbf\{s\}^\{p\},\\mathcal\{H\}^\{p\}=\\mbox\{$\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits$\}\(\\mathbf\{x\}^\{p\},\\mathbf\{v\}^\{p\},\\mathbf\{x\}^\{r\},\\mathbf\{v\}^\{r\}\)โณ\\trianglerightencode pocket into embeddings using๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits
6:for
t=Tt=Tto
11do
7:
\(๐ฑ~0,tg,๐ฑ~0,tg\)=๐๐ผ๐ซ๐ฆ\(๐ฑtg,๐ฏtg,๐ฌp,โp\)\(\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{0,t\},\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{0,t\}\)=\\mbox\{$\\mathop\{\\mathsf\{pcLG\}\}\\limits$\}\(\\mathbf\{x\}^\{g\}\_\{t\},\\mathbf\{v\}^\{g\}\_\{t\},\\mathbf\{s\}^\{p\},\\mathcal\{H\}^\{p\}\)โณ\\trianglerightpredict noise\-free ligand using the๐๐ผ๐ซ๐ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limits
8:
๐ณtโ1๐กโผ๐ฉโ\(0,๐\),๐ฑtโ1g=qโ\(๐ฑtโ1g\|๐ฑtg,๐ฑ~0,tg\)\\mathbf\{z\}^\{\\mathtt\{x\}\}\_\{t\-1\}\\sim\\mathcal\{N\}\(0,\\mathbf\{I\}\),\\quad\\mathbf\{x\}^\{g\}\_\{t\-1\}=q\(\\mathbf\{x\}^\{g\}\_\{t\-1\}\|\\mathbf\{x\}^\{g\}\_\{t\},\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{0,t\}\)โณ\\trianglerightsample๐ฑtโ1g\\mathbf\{x\}^\{g\}\_\{t\-1\}using Gaussian posterior \(Equation[16](https://arxiv.org/html/2607.12349#Sx6.E16)\)
9:
๐ณtโ1๐โผGumbelโ\(0,1\),๐ฏtโ1g=qโ\(๐ฏtโ1g\|๐ฏtg,๐ฑ~0,tg\)\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{t\-1\}\\sim\\text\{Gumbel\}\(0,1\),\\quad\\mathbf\{v\}^\{g\}\_\{t\-1\}=q\(\\mathbf\{v\}^\{g\}\_\{t\-1\}\|\\mathbf\{v\}^\{g\}\_\{t\},\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{0,t\}\)โณ\\trianglerightsample๐ฏtโ1g\\mathbf\{v\}^\{g\}\_\{t\-1\}using categorical posterior \(Equation[17](https://arxiv.org/html/2607.12349#Sx6.E17)\)
10:
ฯ๐ฑโฯ๐ฑโช\{๐ณtโ1๐ก\},ฯ๐ฏโฯ๐ฏโช\{๐ณtโ1๐\}\\tau^\{\\mathbf\{x\}\}\\leftarrow\\tau^\{\\mathbf\{x\}\}\\cup\\\{\\mathbf\{z\}^\{\\mathtt\{x\}\}\_\{t\-1\}\\\},\\;\\;\\tau^\{\\mathbf\{v\}\}\\leftarrow\\tau^\{\\mathbf\{v\}\}\\cup\\\{\\mathbf\{z\}^\{\\mathtt\{v\}\}\_\{t\-1\}\\\}โณ\\trianglerightappend noise to trajectories
11:endfor
12:
๐best=๐ญ๐๐๐๐พ๐ฎ๐๐โ\(ฯ๐ฑ,ฯ๐ฑ,N\)\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\_\{\\text\{best\}\}=\\mathsf\{NoiseOpt\}\(\\tau^\{\\mathbf\{x\}\},\\tau^\{\\mathbf\{x\}\},N\)โณ\\trianglerightAlgorithm[3](https://arxiv.org/html/2607.12349#alg3), noise optimization
13:return
๐best\\mbox\{$\\mathop\{\\mathcal\{D\}\}\\limits$\}\_\{\\text\{best\}\}
The overall๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsframework is summarized in Algorithm[5](https://arxiv.org/html/2607.12349#alg5)\. At a high level,๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsrecords the noise trajectories produced by๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsduring sampling and then optimizes these trajectories to refine the resulting ligand\. Specifically, it first samples the ligand atom count based on the pocket size and initializes atom positions with Gaussian noise and atom types with Gumbel noise\. Conditioned on pocket embeddings produced by๐๐๐ฏ๐ฑ๐ซ\-โ๐พ๐๐ผ\\mathop\{\\mathsf\{\\mbox\{$\\mathop\{\\mathsf\{msPRL\}\}\\limits$\}\}\\text\{\-\}\\mathsf\{enc\}\}\\limits,๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limitsruns the generation process fromt=Tt=Tto11: at each step,๐๐ผ๐ซ๐ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitsestimates the noise\-free ligand\(๐ฑ~0,tg,๐ฑ~0,tg\)\(\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{0,t\},\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{0,t\}\), after which atom positions and types are updated by sampling from the corresponding Gaussian and categorical posteriors\. The injected position and type noises throughout sampling are recorded to the trajectories\(ฯ๐ฑ,ฯ๐ฏ\)\(\\tau^\{\\mathbf\{x\}\},\\tau^\{\\mathbf\{v\}\}\)\. Finally, the noise optimization module refines the sampled ligand by optimizing the trajectories forNNiterations and returns the optimized molecule๐\\mathop\{\\mathcal\{D\}\}\\limits\.
## Data Availability
## Code Availability
## Acknowledgements
This project was made possible, in part, by support from the National Science Foundation grant nos\. 2435819 \(X\.N\.\) and 2450988 \(M\.H\.\), the National Library of Medicine grant no\. 1R01LM014385 \(X\.N\.\), the National Center for Advancing Translational Sciences grant no\. UM1TR004548 \(X\.N\.\), and Sanofi iDEA\-TECH Awards North America \(X\.N\.\)\. Any opinions, findings, and conclusions or recommendations expressed in this manuscript are those of the authors, and do not necessarily reflect the views of the funding agencies\. We thank Benjamin Burns, Reza Averly, Maggie Samaan, and Trieu Nguyen for their help with paper writing\. We are grateful for their careful review, constructive feedback, and assistance in improving the clarity and organization of the manuscript\. We thank Avery Meyer for her contributions to the design and implementation of the graphical user interface, which helps improve the usability and accessibility of the method\.
## Author Contributions
X\.N\. conceived the research and conducted the project administration\. M\.H\. and X\.N\. investigated the research, obtained funding and resources for the research, and supervised the student authors \(X\.N\. supervised R\.G, Z\.C\., and F\.B\.; M\.H\. supervised J\.P\.\)\. R\.G\., Z\.C\. and X\.N\. designed the computational methodologies of๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limits\. J\.P\. and M\.H\. designed the computational methodologies of๐๐บ๐ฎ๐ฏ๐ณ\\mathop\{\\mathsf\{paOPT\}\}\\limits\. R\.G\. conducted data curation, formal analysis, computational methodology \(๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limits\) implementation, result analysis, and visualization\. J\.P\. conducted formal analysis, computational methodology \(๐๐บ๐ฎ๐ฏ๐ณ\\mathop\{\\mathsf\{paOPT\}\}\\limits\) implementation, result analysis, and visualization\. Z\.C\. contributed to the methodology design and implementation \(๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limits\)\. F\.B\. contributed to formal analysis, result analysis, and visualization\. D\.K\. designed the computational chemistry evaluation and experimental validation of the research and contributed to the computational chemistry results analysis and visualization\. J\.K\. and A\.S\. conducted the molecular synthesis experiments and the results analysis\. H\-P\.B\., Y\.L\. and M\.L\. contributed to the computational chemistry evaluation and experimental validation analysis\. L\.I\. conducted biological assays for PD\-L1\. R\.G\., J\.P\., M\.H\., and X\.N\. drafted the original manuscript\. R\.G\., J\.P\., F\.B\., D\.K\., M\.H\., and X\.N\. conducted the manuscript editing and revision\. All authors reviewed the final paper\.
## References
## Appendix AEquivariance and Invariance
### Equivariance
By definition, a functionfโ\(๐\)f\(\\mathsf\{x\}\)is equivariant if all translation and rotation transformations from the special Euclidean group SE\(3\)\[Atz2021\]applied to the input๐โโ3\\mathsf\{x\}\\in\\mathbb\{R\}^\{3\}are mirrored accordingly in the output as follows:
fโ\(๐โ๐\+๐ญ\)=๐โfโ\(๐\)\+๐ญ,f\(\\mathbf\{R\}\\mathsf\{x\}\+\\mathbf\{t\}\)=\\mathbf\{R\}f\(\\mathsf\{x\}\)\+\\mathbf\{t\},\(35\)where,๐ญโโ3\\mathbf\{t\}\\in\\mathbb\{R\}^\{3\}is a translation transformation and๐โโ3ร3\\mathbf\{R\}\\in\\mathbb\{R\}^\{3\\times 3\}\(๐๐ณโ๐=๐\\mathbf\{R\}^\{\\mathsf\{T\}\}\\mathbf\{R\}=\\mathbf\{I\}\) is a rotation transformation\.๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitsleverages GVP and VN\-MLP to ensure๐๐๐ฏ๐ฑ๐ซ\\mathop\{\\mathsf\{msPRL\}\}\\limitsand๐๐ผ๐ซ๐ฆ\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{pcLG\}\}\\limits$\}\}\\limitsare equivariant, effectively capturing the geometric features of objects regardless of any translation or rotation transformations\.
### Invariance
A functionfโ\(๐\)f\(\\mathsf\{x\}\)is invariant if its output remains constant under all translation and rotation transformations of the input๐\\mathsf\{x\}:
fโ\(๐โ๐\+๐ญ\)=fโ\(๐\),f\(\\mathbf\{R\}\\mathsf\{x\}\+\\mathbf\{t\}\)=f\(\\mathsf\{x\}\),\(36\)where๐ญ\\mathbf\{t\}and๐\\mathbf\{R\}represents any translation and rotation transformation, respectively\.๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limitslearns invariant scalar embeddings for pocket atoms and ligand atoms, capturing inherent features \(e\.g\., atom features\) that remain invariant to any translation or rotation transformation\.
## Appendix BForward Diffusion
In the forward process,๐๐ผ๐ซ๐ฆ\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{pcLG\}\}\\limits$\}\}\\limitsadds noises step by step to the atom position \(๐ฑi,tg\\mathbf\{x\}^\{g\}\_\{i,t\}\) and atom feature \(๐ฏi,tg\\mathbf\{v\}^\{g\}\_\{i,t\}\) in the training ligands\. For brevity, in this section, we eliminate the subscriptiiin the notations when no ambiguity arises\. The probability of atom positions๐ฑtg\\mathbf\{x\}^\{g\}\_\{t\}sampled given๐ฑtโ1g\\mathbf\{x\}^\{g\}\_\{t\-1\}, denoted asqโ\(๐ฑtg\|๐ฑtโ1g\)q\(\\mathbf\{x\}^\{g\}\_\{t\}\|\\mathbf\{x\}^\{g\}\_\{t\-1\}\), is defined as follows:
qโ\(๐ฑtg\|๐ฑtโ1g\)=๐ฉโ\(๐ฑtg\|1โฮฒt๐ฑโ๐ฑtโ1g,ฮฒt๐ฑโ๐\),q\(\\mathbf\{x\}^\{g\}\_\{t\}\|\\mathbf\{x\}^\{g\}\_\{t\-1\}\)=\\mathcal\{N\}\(\\mathbf\{x\}^\{g\}\_\{t\}\|\\sqrt\{1\-\\beta^\{\\mathbf\{x\}\}\_\{t\}\}\\mathbf\{x\}^\{g\}\_\{t\-1\},\\beta^\{\\mathbf\{x\}\}\_\{t\}\\mathbf\{I\}\),\(37\)where๐ฉโ\(โ
\)\\mathcal\{N\}\(\\cdot\)is a Gaussian distribution of๐ฑtg\\mathbf\{x\}^\{g\}\_\{t\}with mean1โฮฒt๐ฑโ๐ฑtโ1g\\sqrt\{1\-\\beta\_\{t\}^\{\\mathbf\{x\}\}\}\\mathbf\{x\}^\{g\}\_\{t\-1\}and covarianceฮฒt๐ฑโ๐\\beta\_\{t\}^\{\\mathbf\{x\}\}\\mathbf\{I\}\. The probability of atom features at time steptt,๐ฏtg\\mathbf\{v\}^\{g\}\_\{t\}, given that at time steptโ1t\-1,๐ฏtโ1g\\mathbf\{v\}^\{g\}\_\{t\-1\}is defined as follows:
qโ\(๐ฏtg\|๐ฏtโ1g\)=๐โ\(๐ฏtg\|\(1โฮฒt๐ฏ\)โ๐ฏtโ1g\+ฮฒt๐ฏโ๐/da\),q\(\\mathbf\{v\}^\{g\}\_\{t\}\|\\mathbf\{v\}^\{g\}\_\{t\-1\}\)=\\mathcal\{C\}\(\\mathbf\{v\}^\{g\}\_\{t\}\|\(1\-\\beta^\{\\mathbf\{v\}\}\_\{t\}\)\\mathbf\{v\}^\{g\}\_\{t\-1\}\+\\beta^\{\\mathbf\{v\}\}\_\{t\}\\mathbf\{1\}/d\_\{a\}\),\(38\)where๐\\mathcal\{C\}is a categorical distribution\.
Given the above definitions, the probability of atom positions \(๐ฑtg\\mathbf\{x\}^\{g\}\_\{t\}\) and atom features \(๐ฏtg\\mathbf\{v\}^\{g\}\_\{t\}\) at any time stepttcan be derived from those at the initial time step \(๐ฑ0g\\mathbf\{x\}^\{g\}\_\{0\}and๐ฏ0g\\mathbf\{v\}^\{g\}\_\{0\}\) as follows:
qโ\(๐ฑtg\|๐ฑ0g\)\\displaystyle q\(\\mathbf\{x\}^\{g\}\_\{t\}\|\\mathbf\{x\}^\{g\}\_\{0\}\)=๐ฉโ\(๐ฑtg\|ฮฑยฏt๐ฑโ๐ฑ0g,\(1โฮฑยฏt๐ฑ\)โ๐\),\\displaystyle=\\mathcal\{N\}\(\\mathbf\{x\}^\{g\}\_\{t\}\|\\sqrt\{\\mbox\{$\\mathop\{\\bar\{\\alpha\}\}\\limits$\}^\{\\mathbf\{x\}\}\_\{t\}\}\\mathbf\{x\}^\{g\}\_\{0\},\(1\-\\mbox\{$\\mathop\{\\bar\{\\alpha\}\}\\limits$\}^\{\\mathbf\{x\}\}\_\{t\}\)\\mathbf\{I\}\),\(39\)qโ\(๐ฏtg\|๐ฏ0g\)\\displaystyle q\(\\mathbf\{v\}^\{g\}\_\{t\}\|\\mathbf\{v\}^\{g\}\_\{0\}\)=๐โ\(๐ฏtg\|ฮฑยฏt๐ฏ๐ฏ0g\+\(1โฮฑยฏt๐ฏ\)โ๐/K\),\\displaystyle=\\mathcal\{C\}\(\\mathbf\{v\}^\{g\}\_\{t\}\|\\mbox\{$\\mathop\{\\bar\{\\alpha\}\}\\limits$\}^\{\\mathbf\{v\}\}\_\{t\}\\mathbf\{v\}^\{g\}\_\{0\}\+\(1\-\\mbox\{$\\mathop\{\\bar\{\\alpha\}\}\\limits$\}^\{\\mathbf\{v\}\}\_\{t\}\)\\mathbf\{1\}/K\),\(40\)whereฮฑยฏt๐\\displaystyle\\text\{where \}\\mbox\{$\\mathop\{\\bar\{\\alpha\}\}\\limits$\}^\{\\mathtt\{u\}\}\_\{t\}=โฯ=1tฮฑฯ๐,ฮฑฯ๐=1โฮฒฯ๐,๐=๐ฑโorโ๐ฏ,\\displaystyle=\\displaystyle\\prod\_\{\\tau=1\}^\{t\}\\alpha^\{\\mathtt\{u\}\}\_\{\\tau\},\\ \\alpha^\{\\mathtt\{u\}\}\_\{\\tau\}=1\-\\beta^\{\\mathtt\{u\}\}\_\{\\tau\},\\ \{\\mathtt\{u\}\}=\{\\mathbf\{x\}\}\\text\{ or \}\{\\mathbf\{v\}\},\\;\\;\\;\(41\)whereฮฑยฏt๐\\bar\{\\alpha\}^\{\\mathtt\{u\}\}\_\{t\}is a weight decreasing monotonically from 1 to 0 overt=\[1,T\]t=\[1,T\]\. Specifically,ฮฑยฏt๐\\mbox\{$\\mathop\{\\bar\{\\alpha\}\}\\limits$\}^\{\\mathtt\{u\}\}\_\{t\}\(๐=๐ฑโorโ๐ฏ\\mathtt\{u\}=\{\\mathbf\{x\}\}\\text\{ or \}\{\\mathbf\{v\}\}\) approaches 1 astโ1t\\rightarrow 1, allowing๐ฑtg\\mathbf\{x\}^\{g\}\_\{t\}or๐ฏtg\\mathbf\{v\}^\{g\}\_\{t\}to approximate๐ฑ0g\\mathbf\{x\}^\{g\}\_\{0\}or๐ฏ0g\\mathbf\{v\}^\{g\}\_\{0\}\. Conversely,ฮฑยฏt๐\\mbox\{$\\mathop\{\\bar\{\\alpha\}\}\\limits$\}^\{\\mathtt\{u\}\}\_\{t\}\(๐=๐ฑโorโ๐ฏ\\mathtt\{u\}=\{\\mathbf\{x\}\}\\text\{ or \}\{\\mathbf\{v\}\}\) approaches 0 astโTt\\rightarrow T, which makesqโ\(xTgโฃx0g\)q\(x\_\{T\}^\{g\}\\mid x\_\{0\}^\{g\}\)resemble๐ฉโ\(0,I\)\\mathcal\{N\}\(0,I\)andqโ\(vTgโฃv0g\)q\(v\_\{T\}^\{g\}\\mid v\_\{0\}^\{g\}\)resemble๐โ\(1/da\)\\mathcal\{C\}\(1/d\_\{a\}\)\.
As shown by Ho*et al\.*\[ho2020ddpm\], the ground\-truth Normal posterior of atom positions,pโ\(๐ฑtโ1g\|๐ฑtg,๐ฑ0g\)p\(\\mathbf\{x\}^\{g\}\_\{t\-1\}\|\\mathbf\{x\}^\{g\}\_\{t\},\\mathbf\{x\}^\{g\}\_\{0\}\), could be calculated in a closed form as below:
qโ\(๐ฑtโ1g\|๐ฑtg,๐ฑ0g\)=๐ฉโ\(๐ฑtโ1g\|ฮผโ\(๐ฑtg,๐ฑ0g\),ฮฒ~t๐ฑโ๐\),\\displaystyle q\(\\mathbf\{x\}^\{g\}\_\{t\-1\}\|\\mathbf\{x\}^\{g\}\_\{t\},\\mathbf\{x\}^\{g\}\_\{0\}\)=\\mathcal\{N\}\(\\mathbf\{x\}^\{g\}\_\{t\-1\}\|\\mu\(\\mathbf\{x\}^\{g\}\_\{t\},\\mathbf\{x\}^\{g\}\_\{0\}\),\\tilde\{\\beta\}^\{\\mathbf\{x\}\}\_\{t\}\\mathbf\{I\}\),\(42\)ฮผโ\(๐ฑtg,๐ฑ0g\)=ฮฑยฏtโ1๐ฑโฮฒt๐ฑ1โฮฑยฏt๐ฑโ๐ฑ0g\+ฮฑt๐ฑโ\(1โฮฑยฏtโ1๐ฑ\)1โฮฑยฏt๐ฑโ๐ฑtg,\\displaystyle\\mu\(\\mathbf\{x\}^\{g\}\_\{t\},\\mathbf\{x\}^\{g\}\_\{0\}\)=\\frac\{\\sqrt\{\\bar\{\\alpha\}^\{\\mathbf\{x\}\}\_\{t\-1\}\}\\beta^\{\\mathbf\{x\}\}\_\{t\}\}\{1\-\\bar\{\\alpha\}^\{\\mathbf\{x\}\}\_\{t\}\}\\mathbf\{x\}^\{g\}\_\{0\}\\\!\+\\\!\\frac\{\\sqrt\{\\alpha^\{\\mathbf\{x\}\}\_\{t\}\}\(1\-\\bar\{\\alpha\}^\{\\mathbf\{x\}\}\_\{t\-1\}\)\}\{1\-\\bar\{\\alpha\}^\{\\mathbf\{x\}\}\_\{t\}\}\\mathbf\{x\}^\{g\}\_\{t\},\(43\)ฮฒ~t๐ฑ=1โฮฑยฏtโ1๐ฑ1โฮฑยฏt๐ฑโฮฒt๐ฑ\.\\displaystyle\\tilde\{\\beta\}^\{\\mathbf\{x\}\}\_\{t\}=\\frac\{1\-\\bar\{\\alpha\}^\{\\mathbf\{x\}\}\_\{t\-1\}\}\{1\-\\bar\{\\alpha\}^\{\\mathbf\{x\}\}\_\{t\}\}\\beta^\{\\mathbf\{x\}\}\_\{t\}\.\\;\\;\\;\(44\)Similarly, as shown in Hoogeboom*et al\.*\[hoogeboom22diff\], the ground\-truth categorical posterior of atom featurespโ\(๐ฏtโ1g\|๐ฏtg,๐ฏ0g\)p\(\\mathbf\{v\}^\{g\}\_\{t\-1\}\|\\mathbf\{v\}^\{g\}\_\{t\},\\mathbf\{v\}^\{g\}\_\{0\}\)can be calculated as below:
qโ\(๐ฏtโ1g\|๐ฏtg,๐ฏ0g\)=๐โ\(๐ฏtโ1g\|๐โ\(๐ฏtg,๐ฏ0g\)\),\\displaystyle q\(\\mathbf\{v\}^\{g\}\_\{t\-1\}\|\\mathbf\{v\}^\{g\}\_\{t\},\\mathbf\{v\}^\{g\}\_\{0\}\)=\\mathcal\{C\}\(\\mathbf\{v\}^\{g\}\_\{t\-1\}\|\\mathbf\{c\}\(\\mathbf\{v\}^\{g\}\_\{t\},\\mathbf\{v\}^\{g\}\_\{0\}\)\),\(45\)๐โ\(๐ฏtg,๐ฏ0g\)=๐~/โk=1Kc~k,\\displaystyle\\mathbf\{c\}\(\\mathbf\{v\}^\{g\}\_\{t\},\\mathbf\{v\}^\{g\}\_\{0\}\)=\\tilde\{\\mathbf\{c\}\}/\{\\sum\_\{k=1\}^\{K\}\\tilde\{c\}\_\{k\}\},\(46\)๐~=\[ฮฑt๐ฏโ๐ฏtg\+1โฮฑt๐ฏda\]โ\[ฮฑยฏtโ1๐ฏโ๐ฏ0g\+1โฮฑยฏtโ1๐ฏda\],\\displaystyle\\tilde\{\\mathbf\{c\}\}=\[\\alpha^\{\\mathbf\{v\}\}\_\{t\}\\mathbf\{v\}^\{g\}\_\{t\}\+\\frac\{1\-\\alpha^\{\\mathbf\{v\}\}\_\{t\}\}\{d\_\{a\}\}\]\\odot\[\\bar\{\\alpha\}^\{\\mathbf\{v\}\}\_\{t\-1\}\\mathbf\{v\}^\{g\}\_\{0\}\+\\frac\{1\-\\bar\{\\alpha\}^\{\\mathbf\{v\}\}\_\{t\-1\}\}\{d\_\{a\}\}\],\(47\)where๐โ\(๐ฏtg,๐ฏ0g\)\\mathbf\{c\}\(\\mathbf\{v\}^\{g\}\_\{t\},\\mathbf\{v\}^\{g\}\_\{0\}\)denotes the probability over thedad\_\{a\}classes,c~k\\tilde\{c\}\_\{k\}denotes the likelihood of thekk\-th class, andโ\\odotis the element\-wise product operation\.
## Appendix CBackward Generative Process
In the backward process,๐๐ผ๐ซ๐ฆ\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{pcLG\}\}\\limits$\}\}\\limitsgenerates realistic binding ligands from random noise\. Particularly, conditioned on๐ฑi,tg\\mathbf\{x\}^\{g\}\_\{i,t\}and๐~i,0,tg\\tilde\{\\mathsf\{x\}\}^\{g\}\_\{i,0,t\}, the probabilitypโ\(๐ฑi,tโ1g\|๐ฑi,tg\)p\(\\mathbf\{x\}^\{g\}\_\{i,t\-1\}\|\\mathbf\{x\}^\{g\}\_\{i,t\}\)could be estimated using the approximated posteriorp๐ฏโ\(๐ฑi,tโ1g\|๐ฑi,tg,๐~i,0,tg\)p\_\{\\boldsymbol\{\\Theta\}\}\(\\mathbf\{x\}^\{g\}\_\{i,t\-1\}\|\\mathbf\{x\}^\{g\}\_\{i,t\},\\tilde\{\\mathsf\{x\}\}^\{g\}\_\{i,0,t\}\), as shown in Ho*et al\.*\[ho2020ddpm\]\. Same as Appendix[B](https://arxiv.org/html/2607.12349#A2), we eliminate the subscriptiiin the notations when no ambiguity arises\. The approximation ofpโ\(๐ฑtโ1g\|๐ฑtg\)p\(\\mathbf\{x\}^\{g\}\_\{t\-1\}\|\\mathbf\{x\}^\{g\}\_\{t\}\)is calculated as follows:
pโ\(๐ฑtโ1g\|๐ฑtg\)\\displaystyle p\(\\mathbf\{x\}^\{g\}\_\{t\-1\}\|\\mathbf\{x\}^\{g\}\_\{t\}\)โqโ\(๐ฑtโ1g\|๐ฑtg,๐ฑ~0,tg\)\\displaystyle\\approx q\(\\mathbf\{x\}^\{g\}\_\{t\-1\}\|\\mathbf\{x\}^\{g\}\_\{t\},\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{0,t\}\)\(48\)=๐ฉโ\(๐ฑtโ1g\|ฮผโ\(๐ฑtg,๐ฑ~0,tg\),ฮฒ~t๐ฑโ๐\),\\displaystyle=\\mathcal\{N\}\(\\mathbf\{x\}^\{g\}\_\{t\-1\}\|\\mu\(\\mathbf\{x\}^\{g\}\_\{t\},\\tilde\{\\mathbf\{x\}\}^\{g\}\_\{0,t\}\),\\tilde\{\\beta\}\_\{t\}^\{\\mathbf\{x\}\}\\mathbf\{I\}\),whereฮผโ\(๐ฑtg,๐~0,tg\)\\mu\(\\mathbf\{x\}^\{g\}\_\{t\},\\tilde\{\\mathsf\{x\}\}^\{g\}\_\{0,t\}\)is an estimate ofฮผโ\(๐ฑtg,๐ฑ0g\)\\mu\(\\mathbf\{x\}^\{g\}\_\{t\},\\mathbf\{x\}^\{g\}\_\{0\}\)by replacing๐ฑ0g\\mathbf\{x\}^\{g\}\_\{0\}with its approximation๐~0,tg\\tilde\{\\mathsf\{x\}\}^\{g\}\_\{0,t\}in Equation[42](https://arxiv.org/html/2607.12349#A2.E42)\. Similarly, as shown in Hoogeboom\[hoogeboom22diff\], given๐ฏtg\\mathbf\{v\}^\{g\}\_\{t\}and๐~0,tg\\tilde\{\\mathsf\{v\}\}^\{g\}\_\{0,t\}, the probability of๐ฏtโ1g\\mathbf\{v\}^\{g\}\_\{t\-1\}conditioned on๐ฏtg\\mathbf\{v\}^\{g\}\_\{t\},pโ\(๐ฏtโ1g\|๐ฏtg\)p\(\\mathbf\{v\}^\{g\}\_\{t\-1\}\|\\mathbf\{v\}^\{g\}\_\{t\}\), can be estimated by the approximated posteriorqโ\(๐ฏtโ1g\|๐ฏtg,๐~0,tg\)q\(\\mathbf\{v\}^\{g\}\_\{t\-1\}\|\\mathbf\{v\}^\{g\}\_\{t\},\\tilde\{\\mathsf\{v\}\}^\{g\}\_\{0,t\}\)as below:
pโ\(๐ฏtโ1g\|๐ฏtg\)โqโ\(๐ฏtโ1g\|๐ฏtg,๐~0,tg\)=๐โ\(๐ฏtโ1g\|๐โ\(๐ฏtg,๐~0,tg\)\),\\displaystyle p\(\\mathbf\{v\}^\{g\}\_\{t\-1\}\|\\mathbf\{v\}^\{g\}\_\{t\}\)\\approx q\(\\mathbf\{v\}^\{g\}\_\{t\-1\}\|\\mathbf\{v\}^\{g\}\_\{t\},\\tilde\{\\mathsf\{v\}\}^\{g\}\_\{0,t\}\)=\\mathcal\{C\}\(\\mathbf\{v\}^\{g\}\_\{t\-1\}\|\\mathbf\{c\}\(\\mathbf\{v\}^\{g\}\_\{t\},\\tilde\{\\mathsf\{v\}\}^\{g\}\_\{0,t\}\)\),\(49\)where๐โ\(๐ฏtg,๐~0,tg\)\\mathbf\{c\}\(\\mathbf\{v\}^\{g\}\_\{t\},\\tilde\{\\mathsf\{v\}\}^\{g\}\_\{0,t\}\)is an estimate of๐โ\(๐ฏtg,๐ฏ0g\)\\mathbf\{c\}\(\\mathbf\{v\}^\{g\}\_\{t\},\\mathbf\{v\}^\{g\}\_\{0\}\)by replacing๐ฏ0g\\mathbf\{v\}^\{g\}\_\{0\}with its estimate๐~0,tg\\tilde\{\\mathsf\{v\}\}^\{g\}\_\{0,t\}in Equation[45](https://arxiv.org/html/2607.12349#A2.E45)\.
## Appendix DParameters for Reproducibility
In๐ผ๐๐๐ฃ๐๐๐บ๐\\mathop\{\\mathsf\{conDitar\}\}\\limits, we trained two models๐๐๐ฏ๐ฑ๐ซ\\mathop\{\\mathsf\{msPRL\}\}\\limitsand๐๐ผ๐ซ๐ฆ\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{pcLG\}\}\\limits$\}\}\\limitsfor pocket representation learning and pocket\-conditioned ligand generation, respectively\. We implemented both models using Python 3\.9\.18 and PyTorch 2\.1\.0, together with the corresponding PyTorch Geometric dependencies, including torch\-scatter 2\.1\.2, torch\-cluster 1\.6\.3, and torch\-geometric 2\.6\.1\. The environment is built with CUDA 12\.1 \(pytorch\-cuda 12\.1\)\. We trained both models on an NVIDIA A100 GPU with 40GB memory and a CPU with 80GB memory\.
### Parameters formsPRL\\mathop\{\\mathsf\{msPRL\}\}\\limits
In๐๐๐ฏ๐ฑ๐ซ\\mathop\{\\mathsf\{msPRL\}\}\\limits, we set the dimension of all the hidden layers, including GVP layers \(Equation[3](https://arxiv.org/html/2607.12349#Sx6.E3)and[4](https://arxiv.org/html/2607.12349#Sx6.E4)\) and MLP layers \(Equation[5](https://arxiv.org/html/2607.12349#Sx6.E5)to[9](https://arxiv.org/html/2607.12349#Sx6.E9)\) as 128, and the dimension of both residue scalar and vector embeddings \(๐ฌr\\mathbf\{s\}^\{r\}andโr\\mathcal\{H\}^\{r\}\) and pocket atom scalar and vector embeddings \(๐ฌp\\mathbf\{s\}^\{p\}andโp\\mathcal\{H\}^\{p\}\) as 128\. To represent the pocket structure, we constructed two graphs using thekk\-nearest neighbors based on Euclidean distance: one for pocket atoms withka=16k\_\{a\}=16and one for residues withkr=8k\_\{r\}=8\. We set the layer number of graph neural networks for both atom and residue as 3\. We optimized the๐๐๐ฏ๐ฑ๐ซ\\mathop\{\\mathsf\{msPRL\}\}\\limitsmodel with Adam with its parameters \(0\.950, 0\.999\), learning rate 0\.001, and batch size 64\. We trained๐๐๐ฏ๐ฑ๐ซ\\mathop\{\\mathsf\{msPRL\}\}\\limitsfor maximum 100 epochs and the training tookโผ\\sim36 hours in total\.
### Parameters forpcLG\\mathop\{\\mathsf\{pcLG\}\}\\limits
In๐๐ผ๐ซ๐ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limits, we set the dimension of all the scalar hidden layers, including GVP layers \(Equation[23](https://arxiv.org/html/2607.12349#Sx6.E23),[25](https://arxiv.org/html/2607.12349#Sx6.E25)and[31](https://arxiv.org/html/2607.12349#Sx6.E31)\) and VN\-MLP and MLP layers \(Equation[27](https://arxiv.org/html/2607.12349#Sx6.E27)to[33](https://arxiv.org/html/2607.12349#Sx6.E33)\) as 128\. We set the dimensions of all the vector hidden layers in GVPs as 32\. We set the number of layersLLin๐๐ผ๐ซ๐ฆ\\mathop\{\\mathsf\{pcLG\}\}\\limitsas 10\. We built the distance\-based atomic graphs for ligands using themm\-nearest neighbors based on Euclidean distance withm=8m=8\. For each ligand atom, we consider its nearestna=32n\_\{a\}=32pocket atoms for protein\-ligand interaction learning\. In addition, we consider residue\-level interaction by connecting each ligand atom to itsnr=4n\_\{r\}=4nearest pocket residues\.
In the forward process of the diffusion model, following Guan*et al\.*\[guan2023targetdiff\], we used a sigmoidฮฒ\\betaschedule for the variance scheduleฮฒt๐ฑ\\beta\_\{t\}^\{\\mathbf\{x\}\}of atom positions to add noises into atom positions as below:
ฮฒt๐ฑ=sigmoidโ\(w1โ\(2โt/Tโ1\)\)โ\(w2โw3\)\+w3\\beta\_\{t\}^\{\\mathbf\{x\}\}=\\text\{sigmoid\}\(w\_\{1\}\(2t/T\-1\)\)\(w\_\{2\}\-w\_\{3\}\)\+w\_\{3\}\(50\)in whichwiw\_\{i\}\(ii=1,2, or 3\) withw1=6w\_\{1\}=6,w2=1\.eโ7w\_\{2\}=1\.e\-7andw3=0\.01w\_\{3\}=0\.01are hyperparameters;T=1,000T=1,000is the maximum step\. For atom types, we used a cosineฮฒ\\betaschedule\[nichol2021\]forฮฒt๐ฏ\\beta\_\{t\}^\{\\mathbf\{v\}\}as below:
ฮฑยฏt๐ฏ=fโ\(t\)fโ\(0\),f\(t\)=cos\(t/T\+s1\+sโ
ฯ2\)2\\displaystyle\\bar\{\\alpha\}\_\{t\}^\{\\mathbf\{v\}\}=\\frac\{f\(t\)\}\{f\(0\)\},f\(t\)=\\cos\(\\frac\{t/T\+s\}\{1\+s\}\\cdot\\frac\{\\pi\}\{2\}\)^\{2\}\(51\)ฮฒt๐ฏ=1โฮฑt๐ฏ=1โฮฑยฏt๐ฏฮฑยฏtโ1๐ฏ\\displaystyle\\beta\_\{t\}^\{\\mathbf\{v\}\}=1\-\\alpha\_\{t\}^\{\\mathbf\{v\}\}=1\-\\frac\{\\bar\{\\alpha\}\_\{t\}^\{\\mathbf\{v\}\}\}\{\\bar\{\\alpha\}\_\{t\-1\}^\{\\mathbf\{v\}\}\}in whichssis a hyperparameter and set as 0\.01\. Same as๐๐๐ฏ๐ฑ๐ซ\\mathop\{\\mathsf\{msPRL\}\}\\limits, we optimized๐๐ผ๐ซ๐ฆ\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{pcLG\}\}\\limits$\}\}\\limitsusing Adam with its parameters \(0\.950, 0\.999\), the learning rate 0\.001, and batch size 16\. The training takesโผ\\sim100 hours in total\.
### Parameters forconDitar\-โdev\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits
In๐ผ๐๐๐ฃ๐๐๐บ๐\-โ๐ฝ๐พ๐\\mathop\{\\mbox\{$\\mathop\{\\mathsf\{conDitar\}\}\\limits$\}\\text\{\-\}\{\\mathsf\{dev\}\}\}\\limits, we set the number of optimization iterations toN=10N=10and the step size toฮฑ=0\.1\\alpha=0\.1for noise optimization\. For zero*th*\-order gradient estimation, we useH=4H=4random perturbations with a smoothing parameter ofฮผ=0\.03\\mu=0\.03\.
## Appendix EDetails of๐ข๐ฃ๐ง\\mathop\{\\mathsf\{CDH\}\}\\limitsdataset
Table A1:Summary of๐ข๐ฃ๐ง\\mathop\{\\mathsf\{CDH\}\}\\limitsSimilar Articles
Sesame: Structure-Aware Molecular Generation via Spatial Density-Map Conditioning
This paper introduces Sesame, a diffusion-based molecular generation model that conditions on partial molecular structure and protein pocket via spatial density maps, enabling both de novo generation and fragment-conditioned lead optimization for drug design.
TD3B: Transition-Directed Discrete Diffusion for Allosteric Binder Generation
TD3B is a sequence-based generative framework for designing allosteric binders with specific agonist or antagonist behaviors using transition-directed discrete diffusion. The paper introduces a method to control directional transitions in protein states, addressing limitations of static structure-based design.
From Holo Pockets to Electron Density: GPT-style Drug Design with Density
This paper introduces EDMolGPT, an autoregressive framework that generates 3D molecular conformations from low-resolution electron density point clouds, improving structure-based drug design by leveraging physically meaningful density signals.
Controllable Molecular Generative Foundation Models
Proposes CoMole, a controllable molecular generative foundation model using motif-aware graph diffusion and reinforcement learning, achieving superior controllability across materials and drug discovery benchmarks.
Reading the Cell, Designing the Cure: Perturbation-Conditioned Molecular Diffusion for Function-Oriented Drug Design
This paper formalizes transcriptome-based drug design (TBDD) as a generative inverse problem and proposes CURE, a multi-resolution transcriptome-guided diffusion framework that generates drug molecules conditioned on desired transcriptomic state transitions.