Gene Expression-Informed Jointly Controlled Generative Modeling for Precision Molecular Design
Summary
This paper proposes JoPMol, a jointly controlled precision molecular generative model that integrates gene expression profiles, molecular structure text, and chemical properties to generate personalized drug candidates, outperforming state-of-the-art methods.
View Cached Full Text
Cached at: 07/15/26, 04:17 AM
# Gene Expression–Informed Jointly Controlled Generative Modeling for Precision Molecular Design
Source: [https://arxiv.org/html/2607.11978](https://arxiv.org/html/2607.11978)
Hang Yuan[0009\-0001\-3135\-582X](https://orcid.org/0009-0001-3135-582X)School of Artificial IntelligenceSouth China Normal UniversityFoshanGuangdongChinaD3 CenterThe University of OsakaSuitaOsakaJapanChen Li[0000\-0002\-8784\-8148](https://orcid.org/0000-0002-8784-8148)D3 CenterThe University of OsakaSuitaOsakaJapan,Wenjun Ma[0000\-0003\-4600\-398X](https://orcid.org/0000-0003-4600-398X)School of Artificial IntelligenceSouth China Normal UniversityFoshanGuangdongChinaSchool of Computer ScienceSouth China Normal UniversityGuangzhouGuangdongChina,Tadahiko Murata[0000\-0002\-1654\-3945](https://orcid.org/0000-0002-1654-3945)D3 CenterThe University of OsakaSuitaOsakaJapanandYuncheng Jiang[0000\-0002\-0402\-5382](https://orcid.org/0000-0002-0402-5382)School of Artificial IntelligenceSouth China Normal UniversityFoshanGuangdongChinaSchool of Computer ScienceSouth China Normal UniversityGuangzhouGuangdongChina
###### Abstract\.
Precision molecular design aims to discover personalized drug candidates through joint control of multiple conditions, such as biological relevance and molecular design strategies\. Biological relevance reflects cellular functional states under disease or perturbation conditions, while molecular design strategies provide complementary guidance in terms of structural intentions and property optimization\. In this study, we proposeJoPMol, ajointly controlledprecisionmolecular generative model that integrates biological states encoded by gene expression profiles with molecular structure information expressed in text, and chemical properties quantified by numerical values within a unified modeling framework\. This formulation enables coordinated generation and optimization of candidate molecules under joint condition control\. Experimental results show that JoPMol outperforms state\-of\-the\-art methods across multiple evaluation metrics\. Moreover, JoPMol demonstrates strong generalization ability in both transfer tasks and biologically grounded simulation scenarios, validating its effectiveness for precision molecular design\. The source code is publicly available at[https://github\.com/hala\-yh/JoPMol](https://github.com/hala-yh/JoPMol)\.
## 1\.Introduction
Precision molecular design has attracted increasing attention in personalized drug discovery, aiming to generate candidate molecules that are both biologically relevant and chemically feasible under multiple design constraints\(Zenget al\.,[2022](https://arxiv.org/html/2607.11978#bib.bib42)\)\. In practice, the effectiveness of a drug candidate depends on several complementary factors, including its compatibility with biological systems\(Li and Yamanishi,[2025a](https://arxiv.org/html/2607.11978#bib.bib25); Kaitoh and Yamanishi,[2021](https://arxiv.org/html/2607.11978#bib.bib7)\), its structural information\(Edwardset al\.,[2022](https://arxiv.org/html/2607.11978#bib.bib23); Gonget al\.,[2024](https://arxiv.org/html/2607.11978#bib.bib21)\), and its chemical properties\(Pathaket al\.,[2025](https://arxiv.org/html/2607.11978#bib.bib45); Inukaiet al\.,[2025](https://arxiv.org/html/2607.11978#bib.bib16)\)\. Capturing these diverse aspects typically requires information from multiple heterogeneous sources that describe different facets of the drug discovery process\(Zhouet al\.,[2025](https://arxiv.org/html/2607.11978#bib.bib46)\)\. Effectively integrating such heterogeneous information remains a challenge for precision molecular design\.
From a biological perspective, gene expression profiles provide valuable information about disease states and cellular responses to drug perturbations, offering biologically grounded guidance for precision molecular design\(Matsukiyoet al\.,[2024](https://arxiv.org/html/2607.11978#bib.bib41); Loeffleret al\.,[2024](https://arxiv.org/html/2607.11978#bib.bib38)\)\. These genetic signatures capture complex cellular phenotypes and have increasingly been used to inform drug discovery and molecular design\(Kaitoh and Yamanishi,[2021](https://arxiv.org/html/2607.11978#bib.bib7)\)\. However, gene expression profiles inherently exhibit high dimensionality, substantial noise, and complex inter\-gene dependencies without explicit structural correspondence to molecules\(Li and Yamanishi,[2025a](https://arxiv.org/html/2607.11978#bib.bib25)\)\. Moreover, expression patterns vary significantly across disease types and individual patients, further increasing variability in the underlying biological signals\(Kanget al\.,[2025](https://arxiv.org/html/2607.11978#bib.bib40)\)\. These characteristics make it challenging to extract reliable and structurally relevant features for guiding molecular design when relying solely on gene expression profiles\.
Beyond biological context, precision molecular design also needs to consider structural design intentions\(Wanget al\.,[2019](https://arxiv.org/html/2607.11978#bib.bib54)\)and chemical property optimization\(Liuet al\.,[2018](https://arxiv.org/html/2607.11978#bib.bib56)\)\. One line of research approaches this problem through text\-driven molecular design, where natural language descriptions are used to express structural requirements such as functional groups or molecular structures\(Luoet al\.,[2024](https://arxiv.org/html/2607.11978#bib.bib37); Jinet al\.,[2018](https://arxiv.org/html/2607.11978#bib.bib55)\)\. These descriptions provide an interpretable interface that allows researchers to communicate design intentions in a flexible and human\-readable manner\(Wenget al\.,[2025](https://arxiv.org/html/2607.11978#bib.bib36)\)\. However, structural design intentions are typically defined at the molecular level and lack explicit alignment with disease\-specific biological contexts, limiting their ability to provide precise guidance for therapeutic objectives\(Edwardset al\.,[2022](https://arxiv.org/html/2607.11978#bib.bib23)\)\. Consequently, text\-driven approaches tend to focus on satisfying structural constraints, while overlooking the optimization of biologically relevant and pharmacologically important properties\. Quantitative properties such as solubility, synthesizability, and drug\-likeness provide measurable optimization objectives that encourage generated molecules to satisfy practical design criteria\(Khateret al\.,[2025](https://arxiv.org/html/2607.11978#bib.bib29)\)\. Property\-guided approaches therefore enable models to explore chemically meaningful regions of the molecular design space\(Konget al\.,[2024](https://arxiv.org/html/2607.11978#bib.bib31); Igashovet al\.,[2024](https://arxiv.org/html/2607.11978#bib.bib2)\)\. Nevertheless, optimizing numerical properties alone often focuses on theoretical objectives defined in chemical space, without incorporating biological information from real cellular systems\. As a result, designed molecules may satisfy predefined property constraints but remain disconnected from the biological contexts in which drugs ultimately function\.
These signals indicate the importance of integrating multiple sources of information for precision molecular design\. However, most existing molecular design frameworks rely on a single source of information or only partially incorporate multiple design conditions\(Moretet al\.,[2020](https://arxiv.org/html/2607.11978#bib.bib48)\)\. Biological guidance, structural design intentions, and chemical properties originate from heterogeneous modalities with distinct semantic roles and representation spaces, and are therefore often treated in isolation\. Such practices fail to capture their intrinsic dependencies and complementary effects, leading to suboptimal control over molecular design\(Wanget al\.,[2026](https://arxiv.org/html/2607.11978#bib.bib49)\)\. In particular, to the best of our knowledge, there is currently no molecular design framework that jointly integrates gene expression profiles, structural design intentions, and chemical properties within a unified generative model\. Two key challenges contribute to this limitation\. First, these modalities are typically collected from independent data sources, lacking unified and well\-aligned datasets that jointly capture biological signals, structural intentions, and chemical properties\. Second, they differ fundamentally in representation spaces and semantic granularities, which makes their coordinated modeling within a unified generative framework inherently challenging\.
To address these limitations, we proposeJoPMol, ajointly controlledprecisionmolecular generative model for precision molecular design\. JoPMol jointly conditions molecular design on three complementary sources of information: biological states encoded by gene expression profiles, structural design intentions expressed in text, and chemical optimization objectives specified by numerical property values\. To effectively integrate these heterogeneous signals, JoPMol employs source\-specific feature extractors to learn modality\-aware representations and introduces a bidirectional cross\-attention mechanism to enable cross\-modal interaction and information fusion\. The resulting unified representation is used to guide a diffusion\-based molecular generator, enabling end\-to\-end molecular design and optimization under multiple design constraints\. The main contributions are summarized as follows:
- •Jointly controlled precision molecular design: We introduce a jointly controlled generation approach under a multi\-condition setting, enabling precision molecular design that simultaneously accounts for biological relevance, structural design intentions, and property optimization targets\.
- •Integration of design and optimization objectives: We propose JoPMol, which learns latent representations of heterogeneous control signals and incorporates chemical properties as optimization objectives, enabling jointly controlled molecular design and optimization in a unified framework\.
- •Superior performance for precision molecular design: JoPMol consistently outperforms state\-of\-the\-art \(SOTA\) baselines across multiple evaluation metrics and diverse experimental settings, demonstrating its effectiveness and robustness for precision molecular design\.
## 2\.Related Works
### 2\.1\.Gene\-Guided Molecular Design
Gene expression profiles characterize cellular responses to molecular perturbations, providing biologically relevant signals for molecular design\(Matsukiyoet al\.,[2024](https://arxiv.org/html/2607.11978#bib.bib41)\)\. These profiles provide a biologically grounded representation of disease states, making them a promising yet challenging signal for guiding molecular design in practice\. TRIOMPHE\(Kaitoh and Yamanishi,[2021](https://arxiv.org/html/2607.11978#bib.bib7)\)retrieves proxy molecules associated with gene expression profiles that are most similar to a target profile and uses them as inputs to a variational autoencoder \(VAE\)\(Pinheiro Cinelliet al\.,[2021](https://arxiv.org/html/2607.11978#bib.bib27)\)to generate molecules\. Such profiles do not directly participate in the generation process, and similarity\-based retrieval provides only a coarse approximation of complex biological conditions\. In contrast, Gx2Mol\(Li and Yamanishi,[2025b](https://arxiv.org/html/2607.11978#bib.bib26)\)establishes an end\-to\-end mapping between genes and molecular structures, while GxVAEs\(Li and Yamanishi,[2025a](https://arxiv.org/html/2607.11978#bib.bib25)\)further extend this line of work by jointly modeling gene expression profiles and inducing molecules using dual VAEs\. Although these methods demonstrate the feasibility of leveraging biological signals for molecular design, they still treat gene expression profiles as the sole conditioning input and lack explicit mechanisms to jointly integrate structural design intentions and chemical property optimization within a unified generative framework\(Guan and Wang,[2024](https://arxiv.org/html/2607.11978#bib.bib24)\)\. As a result, their ability to support coordinated optimization under joint condition control remains limited in precision molecular design\.
Figure 1\.Overview of the jointly controlled molecular generation framework for precision molecular design\. \(a\)Joint condition control:Three jointly controlled conditions are extracted, including gene expression profiles, textual structural intentions, and numerical property values\. These complementary signals characterize biological relevance and molecular design strategies\. \(b\)Multi\-condition learning:Biological and structural representations are used, and then the resulting features are further integrated with chemical property embeddings to form a joint conditional representation\. \(c\)Conditioned diffusion generation\.The joint control embedding𝐳c\\mathbf\{z\}\_\{c\}guides the reverse diffusion process to progressively refine molecular representations, enabling precise molecular design and property optimization\.
### 2\.2\.Text\-Based Molecular Design
Text\-based structural intentions provide an interpretable interface for specifying structural patterns of molecules, enabling controllable molecular generation beyond purely data\-driven representations in a flexible and human\-interpretable manner\(Denget al\.,[2025](https://arxiv.org/html/2607.11978#bib.bib20)\)\. MolT5\(Edwardset al\.,[2022](https://arxiv.org/html/2607.11978#bib.bib23)\)adopts a Transformer\-based architecture to map natural language descriptions into molecular sequences\. Notably, its autoregressive decoding paradigm often suffers from error sequence sorting and limited controllability over fine\-grained structural modeling\. Text\+ChemT5\(Christofidelliset al\.,[2023](https://arxiv.org/html/2607.11978#bib.bib22)\)enhances modeling capability through a multitask framework that jointly integrates chemical and natural language representations, but it remains constrained by autoregressive generation, making it difficult to balance semantic fidelity and structural consistency\. TGM\-DLM\(Gonget al\.,[2024](https://arxiv.org/html/2607.11978#bib.bib21)\)adopts a diffusion\-based framework to avoid the limitations of autoregressive generation\. It realizes text\-based molecular design through iterative denoising and conditional guidance, and further applies a nearest\-neighbor projection to map continuous latent representations onto discrete molecular tokens, enabling the recovery of molecular sequences\. Such methods show the feasibility of using natural language descriptions to express structural design intentions for precision molecular design\. However, most existing text\-guided models do not explicitly coordinate structural intentions with physicochemical property optimization or biological context, which limits their ability to support coherent and stable control under multiple design objectives\.
### 2\.3\.Property\-Oriented Molecular Design
Explicit numerical property values provide direct and quantitative optimization targets for molecular design, enabling fine\-grained control over properties in a controllable and interpretable manner\(Hou,[2025](https://arxiv.org/html/2607.11978#bib.bib15)\)\. MolSearch\(Sunet al\.,[2022](https://arxiv.org/html/2607.11978#bib.bib18)\)formulates the generation as a neural search process, iteratively producing and scoring candidate molecules to refine property optimization trajectories, but relies on heuristic scoring and does not preserve molecular semantic structures during optimization\. DyMol\(Shinet al\.,[2024](https://arxiv.org/html/2607.11978#bib.bib17)\)adopts a reinforcement learning algorithm that updates policies based on oracle feedback to optimize properties under conditions, while the learning process remains sensitive to reward design and exhibits limited structural controllability\. FRATTVAE\(Inukaiet al\.,[2025](https://arxiv.org/html/2607.11978#bib.bib16)\)incorporates structure\-aware attention into VAEs to improve property\-controllable design\. Nevertheless, property constraints are not aligned with structural semantics, which limits fine\-grained structural control\(Angeloet al\.,[2023](https://arxiv.org/html/2607.11978#bib.bib14)\)\. Moreover, such methods focus on single properties or limited property combinations and struggle to effectively coordinate with complex biological signals or design strategies, limiting their application in multi\-condition molecular design\. In this study, we establish jointly controlled conditions by associating gene expression profiles with structural design intentions, and incorporate numerical property values as explicit optimization objectives into the precision molecular design framework\.
## 3\.Methodology
### 3\.1\.Problem Definition
Precision molecular design aims to generate candidate molecules that satisfy multiple objectives under heterogeneous conditions, including biological relevance, structural intentions, and chemical property requirements\. We formalize these heterogeneous design conditions as three types of inputs\. Specifically, biological relevance is represented by a gene expression profile𝐆=\[g1,g2,…,gK\]∈ℝK\\mathbf\{G\}=\[g\_\{1\},g\_\{2\},\\ldots,g\_\{K\}\]\\in\\mathbb\{R\}^\{K\}, whereKKdenotes the number of genes andgkg\_\{k\}represents the expression value of thekk\-th gene\. Structural intentions are represented as a textual sequence𝐒=\[s1,s2,…,sN\]\\mathbf\{S\}=\[s\_\{1\},s\_\{2\},\\ldots,s\_\{N\}\], whereNNdenotes the sequence length and eachsns\_\{n\}corresponds to a token describing molecular design intent\. Chemical property requirements are represented as a property vector𝐏=\[p1,p2,…,pI\]∈ℝI\\mathbf\{P\}=\[p\_\{1\},p\_\{2\},\\ldots,p\_\{I\}\]\\in\\mathbb\{R\}^\{I\}, whereIIdenotes the number of selected properties andpip\_\{i\}represents a numerical chemical property\. The target is a molecular sequence𝐗=\[x1,x2,…,xM\]\\mathbf\{X\}=\[x\_\{1\},x\_\{2\},\\ldots,x\_\{M\}\], represented in SELFIES, whereMMdenotes the maximum sequence length andxmx\_\{m\}is a discrete token\. The objective is to learn a conditional generative modelpθ\(𝐗∣𝐆,𝐒,𝐏\)p\_\{\\theta\}\(\\mathbf\{X\}\\mid\\mathbf\{G\},\\mathbf\{S\},\\mathbf\{P\}\), which enables controllable molecular generation under jointly conditioned heterogeneous modalities, including gene expression profiles, structural design intentions, and chemical property constraints\.
Theory Motivation: Recent studies in complex systems and representation learning suggest that effective modeling of complex phenomena requires decomposing the system into modules with distinct functional roles\(Battagliaet al\.,[2018](https://arxiv.org/html/2607.11978#bib.bib51); Goyal and Bengio,[2022](https://arxiv.org/html/2607.11978#bib.bib50)\)\. Instead of forcing heterogeneous signals into a unified representation, a modular and compositional perspective advocates modeling different factors separately according to their roles in the generation process\. We adopt this perspective for precision molecular design\. Gene expression profiles𝐆\\mathbf\{G\}and structural design intent𝐒\\mathbf\{S\}provide rich semantic information that constrains the space of feasible molecular structures by specifying biological context and structural requirements\. In contrast, chemical properties𝐏\\mathbf\{P\}exhibit a many\-to\-one relationship with molecular structures, where diverse molecules can share similar property values\. As a result, property signals alone are insufficient to determine structure, and instead act as constraints that bias the generation toward desired objectives\. This functional distinction leads to a structured formulation:
\(1\)p\(𝐗∣𝐆,𝐒,𝐏\)=p\(𝐗∣𝐙sem,𝐏\),𝐙sem=f\(𝐆,𝐒\),p\(\\mathbf\{X\}\\mid\\mathbf\{G\},\\mathbf\{S\},\\mathbf\{P\}\)=p\(\\mathbf\{X\}\\mid\\mathbf\{Z\}\_\{\\text\{sem\}\},\\mathbf\{P\}\),\\quad\\mathbf\{Z\}\_\{\\text\{sem\}\}=f\(\\mathbf\{G\},\\mathbf\{S\}\),wheref\(⋅\)f\(\\cdot\)denotes a learnable mapping that integrates gene expression profiles and structural intent into a unified semantic representation, and𝐙sem\\mathbf\{Z\}\_\{\\text\{sem\}\}captures the resulting structural semantics under biological context\. The property vector𝐏\\mathbf\{P\}modulates the generation process as a separate conditioning signal without directly determining the structure\.
Following this observation, JoPMol adopts a hierarchical multimodal fusion strategy, where the aligned features of gene expression profiles and structural design intentions𝐙sem\\mathbf\{Z\}\_\{\\text\{sem\}\}are first integrated to construct semantic information, and property signals are introduced as optimization objectives to guide the conditional diffusion process, enabling jointly controlled molecular design under heterogeneous conditions\. Figure[1](https://arxiv.org/html/2607.11978#S2.F1)illustrates the JoPMol framework designed based on this principle, consisting of three key components\. The joint condition control module parses and organizes jointly controlled conditions\. The multi\-condition learning module performs representation learning over multiple control signals and constructs unified conditional representations\. The conditioned diffusion generation module enables precise molecular design and property optimization under joint control\.
### 3\.2\.Joint Condition Modeling
We formulate molecular generation under joint control by constructing a unified condition tuple𝒞=\(𝐆,𝐒,𝐏,𝐗\)\\mathcal\{C\}=\(\\mathbf\{G\},\\mathbf\{S\},\\mathbf\{P\},\\mathbf\{X\}\), which integrates biological response signals, structural design intent, and property optimization targets\. Given an induced molecule, its perturbation elicits transcriptional responses in cells, resulting in a differential gene expression profile𝐆\\mathbf\{G\}that reflects the underlying biological state\. The molecule itself is represented as a SELFIES sequence𝐗\\mathbf\{X\}\(Krennet al\.,[2020](https://arxiv.org/html/2607.11978#bib.bib13)\)\. To associate biological signals with explicit design strategies, we derive structural intentions and optimization objectives from the molecule\. A structural design intent𝐒\\mathbf\{S\}is generated using BioT5\(Peiet al\.,[2023](https://arxiv.org/html/2607.11978#bib.bib12)\), adapted using the PubChem database\(Kimet al\.,[2016](https://arxiv.org/html/2607.11978#bib.bib57)\), to express structural patterns and drug\-like characteristics\. All generated structural descriptions are manually reviewed, with detailed annotation procedures provided in Appendix[B](https://arxiv.org/html/2607.11978#A2)\. In addition, a set of quantitative chemical properties𝐏\\mathbf\{P\}is computed to provide interpretable optimization targets for molecular design\. This joint condition formulation establishes a coherent connection between biological response, structural semantics, and property constraints, enabling controlled molecular design under heterogeneous conditions\.
### 3\.3\.Multi\-Condition Learning
We realize the𝐙sem=f\(𝐆,𝐒\)\\mathbf\{Z\}\_\{\\text\{sem\}\}=f\(\\mathbf\{G\},\\mathbf\{S\}\)by learning a unified representation that integrates biological relevant and structural intent, while treating chemical properties𝐏\\mathbf\{P\}as separate conditioning signals\.
Biological Representation Learning\.Gene expression profiles exhibit high dimensionality and substantial noise, making it challenging to directly employ them as conditioning signals for generative modeling\(Li and Yamanishi,[2025a](https://arxiv.org/html/2607.11978#bib.bib25)\)\. We adopt a VAE to learn a compact and smooth biological latent representation\. The VAE encodes𝐆\\mathbf\{G\}into a latent variable𝐳g\\mathbf\{z\}\_\{g\}, which captures essential transcriptomic patterns while filtering out spurious variations\. By enforcing a continuous and approximately Gaussian latent manifold, the VAE provides well\-behaved conditioning signals that are compatible with diffusion\-based generative modeling and improve sampling stability under heterogeneous biological conditions\([Venkatramanet al\.,](https://arxiv.org/html/2607.11978#bib.bib44)\)\. The VAE encoder is trained by minimizing the evidence lower bound:
\(2\)ℒG=𝔼qϕ\(𝐳g∣𝐆\)\[‖𝐆^−𝐆‖22\]\+βKL\(qϕ\(𝐳g∣𝐆\)∥p\(𝐳g\)\),\\mathcal\{L\}\_\{\\text\{G\}\}=\\mathbb\{E\}\_\{q\_\{\\phi\}\(\\mathbf\{z\}\_\{g\}\\mid\\mathbf\{G\}\)\}\\left\[\\left\\\|\\hat\{\\mathbf\{G\}\}\-\\mathbf\{G\}\\right\\\|\_\{2\}^\{2\}\\right\]\+\\beta\\,\\mathrm\{KL\}\\left\(q\_\{\\phi\}\(\\mathbf\{z\}\_\{g\}\\mid\\mathbf\{G\}\)\\,\\\|\\,p\(\\mathbf\{z\}\_\{g\}\)\\right\),whereqϕ\(⋅\)q\_\{\\phi\}\(\\cdot\)denotes the encoder posterior distribution,𝐆^\\hat\{\\mathbf\{G\}\}is the reconstructed gene expression,p\(⋅\)p\(\\cdot\)is the prior distribution, andβ\\betacontrols the strength of latent regularization\. The learned latent variable𝐳g\\mathbf\{z\}\_\{g\}provides a continuous and biological embedding space that facilitates interpolation and generalization across heterogeneous perturbation conditions\. For compatibility with token\-based attention mechanisms,𝐳g\\mathbf\{z\}\_\{g\}is further reshaped and projected into a token\-level biological embedding𝐙g∈ℝTg×D\\mathbf\{Z\}\_\{g\}\\in\\mathbb\{R\}^\{T\_\{g\}\\times D\}, whereTgT\_\{g\}represents the number of tokens, andDDdenotes the embedding dimension\.
Structural Design Intention Representation Learning\.Structural design intentions are derived from the induced molecule and expressed as semantic descriptions that encode molecular structures\. We employ BioLink\-BERT\(Yasunagaet al\.,[2022](https://arxiv.org/html/2607.11978#bib.bib11)\)to transform the structural semantic representation𝐒\\mathbf\{S\}into a contextualized embedding𝐙s=fS\(𝐒\)∈ℝTs×D\\mathbf\{Z\}\_\{s\}=f\_\{S\}\(\\mathbf\{S\}\)\\in\\mathbb\{R\}^\{T\_\{s\}\\times D\}, wherefS\(⋅\)f\_\{S\}\(\\cdot\)denotes the BioLink\-BERT encoder,TsT\_\{s\}represents the number of structural tokens\. The resulting embedding𝐙s\\mathbf\{Z\}\_\{s\}provides interpretable structural guidance that complements the latent biological representation\.
Biological–Structural Semantic Alignment\.To construct the unified semantic representation𝐙sem=f\(𝐆,𝐒\)\\mathbf\{Z\}\_\{\\text\{sem\}\}=f\(\\mathbf\{G\},\\mathbf\{S\}\), we employ a bi\-directional cross\-attention mechanism that enables reciprocal information exchange between𝐙g\\mathbf\{Z\}\_\{g\}and𝐙s\\mathbf\{Z\}\_\{s\}\. Specifically, biological features attend to structural semantics as𝐙g′=𝐙g\+𝐂𝐀\(𝐙g,𝐙s,𝐙s\)\\mathbf\{Z\}\_\{g\}^\{\\prime\}=\\mathbf\{Z\}\_\{g\}\+\\mathbf\{CA\}\(\\mathbf\{Z\}\_\{g\},\\mathbf\{Z\}\_\{s\},\\mathbf\{Z\}\_\{s\}\), allowing gene\-level representations to absorb structurally relevant information, where𝐂𝐀\(⋅\)\\mathbf\{CA\}\(\\cdot\)denotes the multi\-head cross\-attention operator\. Structural features attend to biological representations as𝐙s′=𝐙s\+𝐂𝐀\(𝐙s,𝐙g,𝐙g\)\\mathbf\{Z\}\_\{s\}^\{\\prime\}=\\mathbf\{Z\}\_\{s\}\+\\mathbf\{CA\}\(\\mathbf\{Z\}\_\{s\},\\mathbf\{Z\}\_\{g\},\\mathbf\{Z\}\_\{g\}\), enabling structural tokens to become aware of underlying biological contexts\. This reciprocal interaction promotes mutual calibration between conditions\.
Joint Conditional Representation Construction\.Following the distinction between structural infomation and property constraints, we decouple semantic representation learning from property conditioning\. After semantic alignment, biological and structural features are aggregated and normalized to form a unified semantic context\. Rather than adopting an early\-fusion strategy that entangles heterogeneous modalities, we preserve the semantic hierarchy by isolating structural representation from quantitative objective control\. Chemical properties𝐏\\mathbf\{P\}are projected into the latent space via a linear mappingfproj\(⋅\)f\_\{\\text\{proj\}\}\(\\cdot\)to produce property control tokens, which are incorporated as prefix conditioning signals\. This design enables explicit and stable regulation of molecular properties without disrupting the underlying semantic structure\. The joint conditional representation is constructed as:
\(3\)𝐙c=\[fproj\(𝐏\)⊕LN\(𝐙g′⊕𝐙s′\)\],\\mathbf\{Z\}\_\{\\text\{c\}\}=\\Big\[f\_\{\\text\{proj\}\}\(\\mathbf\{P\}\)\\;\\oplus\\;\\mathrm\{LN\}\\big\(\\mathbf\{Z\}\_\{g\}^\{\\prime\}\\oplus\\mathbf\{Z\}\_\{s\}^\{\\prime\}\\big\)\\Big\],where⊕\\oplusdenotes concatenation along the token dimension andLN\(⋅\)\\mathrm\{LN\}\(\\cdot\)denotes layer normalization\. The resulting𝐙c\\mathbf\{Z\}\_\{\\text\{c\}\}serves as the unified conditioning representation for the subsequent diffusion\-based molecular generator, enabling integrated design and property\-oriented optimization under jointly controlled conditions\.
### 3\.4\.Jointly Controlled Generative Modeling
We embed the SELFIES sequence𝐗\\mathbf\{X\}of the molecule into a continuous latent space and perform diffusion on the resulting embeddings, enabling modeling of discrete molecular sequences\. Specifically, the tokens are mapped into embeddings𝐄0=Emb\(𝐗\)∈ℝL×D\\mathbf\{E\}\_\{0\}=\\mathrm\{Emb\}\(\\mathbf\{X\}\)\\in\\mathbb\{R\}^\{L\\times D\}, whereLLdenotes the sequence length\. Following the standard forward diffusion process, Gaussian noise is integrated as
\(4\)𝐄t=α¯t𝐄0\+1−α¯tϵ,ϵ∼𝒩\(𝟎,𝐈\),\\mathbf\{E\}\_\{t\}=\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\mathbf\{E\}\_\{0\}\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\\boldsymbol\{\\epsilon\},\\quad\\boldsymbol\{\\epsilon\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\),whereα¯t\\bar\{\\alpha\}\_\{t\}denotes the cumulative noise schedule\.
JoPMol is trained to learn a jointly conditioned denoising networkϵθ\(𝐄t,t;𝐙c\)\\epsilon\_\{\\theta\}\(\\mathbf\{E\}\_\{t\},t;\\mathbf\{Z\}\_\{\\text\{c\}\}\)to predict the noise, conditioned on the joint control representation𝐙c\\mathbf\{Z\}\_\{\\text\{c\}\}and the diffusion timesteptt\. Let𝐄^0\\hat\{\\mathbf\{E\}\}\_\{0\}be the recovered clean embedding\. The overall jointly controlled training objective is formulated as
\(5\)ℒ=𝔼t,ϵ\[‖ϵ−ϵθ\(𝐄t,t;𝐙c\)‖22\+ℒcon\(𝐄^0,𝐗\)\+ℒT\+λℒcos\],\\mathcal\{L\}=\\mathbb\{E\}\_\{t,\\boldsymbol\{\\epsilon\}\}\\Big\[\\\|\\boldsymbol\{\\epsilon\}\-\\epsilon\_\{\\theta\}\(\\mathbf\{E\}\_\{t\},t;\\mathbf\{Z\}\_\{\\text\{c\}\}\)\\\|\_\{2\}^\{2\}\+\\mathcal\{L\}\_\{\\text\{con\}\}\(\\hat\{\\mathbf\{E\}\}\_\{0\},\\mathbf\{X\}\)\+\\mathcal\{L\}\_\{T\}\+\\lambda\\mathcal\{L\}\_\{\\text\{cos\}\}\\Big\],where the first term corresponds to the standard denoising objective in diffusion models\.ℒcon\\mathcal\{L\}\_\{\\text\{con\}\}enforces discrete token consistency via vocabulary logits, andℒT\\mathcal\{L\}\_\{T\}stabilizes the terminal diffusion We further introduce a cosine alignment lossℒcos\\mathcal\{L\}\_\{\\text\{cos\}\}to encourage semantic consistency between the biological\-related and structural\-related representations in the shared semantic space\. The weighting coefficientλ\\lambdacontrols the contribution of the alignment objective, balancing biological and structural signals in the model\. Att=0t=0, we match the predicted embedding with the original embedding to further stabilize discrete reconstruction in JoPMol\.
## 4\.Experiments
### 4\.1\.Experimental Setup
Datasets\.Three cancer\-related datasets were used\. Each dataset consists of a 978\-dimensional gene expression profile, the corresponding molecule represented in SELFIES format, structural design intentions, and numerical chemical property values\. Table[1](https://arxiv.org/html/2607.11978#S4.T1)reports dataset statistics, with more details in Appendix[A](https://arxiv.org/html/2607.11978#A1)\.
Table 1\.Statistics of the datasets\. MLen: average SELFIES length; GDim: gene expression dimensionality; TLen: average structural text length; MW: average molecular weight\.Table 2\.Comparison of JoPMol with baseline methods across structure consistency, sequence similarity, property optimization, and overall statistics\. Best results are highlighted in bold, and second\-best results are underlined\.- •TrainSet and TestSetare collected from a human breast cancer cell line under diverse chemical perturbations, comprising 13,755 small molecules\(Duanet al\.,[2014](https://arxiv.org/html/2607.11978#bib.bib10)\), randomly split at a ratio of 9:1 for training and evaluation of JoPMol\.
- •TransferSetconstructed using data from the same breast cancer cell line under genetic perturbations, including gene overexpression \(AKT1, AKT2, AURKB, CTSK, EGFR, HDAC1, MTOR, and PIK3CA\) and gene knockdown \(SMAD3 and TP53\), used to analyze the structural consistency between designed molecules and target\-associated ligands in transfer tasks\.
- •DiseaseSetderived from the CREEDS database\(Wanget al\.,[2016](https://arxiv.org/html/2607.11978#bib.bib9)\), constructed by averaging gene expression profiles from multiple gastric cancer patients, used as a disease\-level case study for evaluating model performance in biologically grounded and clinically relevant simulation scenarios\.
Evaluation Measures\.We evaluate the quality of the designed molecules from four complementary perspectives, including molecular structure, molecular sequence representation, chemical properties, and statistical metrics\.
- •Structure consistency:ECFP\(extended\-connectivity fingerprints\) is used to quantify the consistency of local substructures between designed molecules and the targets\.APFP\(atom pair fingerprints\) encodes atom\-type pairs together with their topological distances, effectively capturing global structural characteristics\.Diceis employed to compute the similarity between two APFP representations for structural comparison\.
- •Sequence similarity:BLEUmeasures the n\-gram overlap between produced SELFIES sequences and the target sequences\.Leven\(Levenshtein\) distance quantifies the minimum number of edit operations required to transform a generated SELFIES into the corresponding target sequence, whileLCS\(longest common subsequence\) evaluates overall sequential consistency\.
- •Property optimization:LogPcharacterizes molecular lipophilicity,SAreflects synthetic accessibility, andQEDmeasures overall drug\-likeness\. These properties are selected as they capture complementary aspects of molecular quality, including pharmacokinetic behavior, synthetic feasibility, and overall drug\-likeness, and are widely adopted in molecular generation studies\. Greater agreement between the average property values of designed molecules and the corresponding test\-set statistics indicates improved distributional consistency and controllability\.
- •Overall statistics:Diversitymeasures structural diversity among candidate molecules\.Totalscore integrates validity, uniqueness, and novelty, providing an overall measure of molecular design performance\.FCD\(Fréchet ChemNet distance\) quantifies the distributional similarity between designed molecules and the targets in the learned chemical feature space\.Rankis obtained by averaging each model’s ranking positions across all metrics, where a lower rank indicates better overall performance\. Additional details are provided in Appendix[D](https://arxiv.org/html/2607.11978#A4)\.
Implementation Details\.JoPMol is implemented in PyTorch\. The gene encoder adopts multi\-layer feedforward architectures with a final latent embedding dimension of 64\. Structural intentions are encoded into fixed\-length semantic representations, while SELFIES tokens are embedded into a 32\-dimensional trainable embedding space with a maximum sequence length of 256\. The joint conditional diffusion is trained for 2,000 timesteps using a cosine noise schedule and optimized with AdamW at a learning rate of1×10−41\\times 10^\{\-4\}and a dropout rate of 0\.1 for regularization\. A uniform skip\-sampling strategy is adopted during reverse diffusion to improve sampling efficiency and generation stability\. We analyze the sensitivity of the weighting coefficientλ\\lambdain Table[E\.1](https://arxiv.org/html/2607.11978#A5.T1)\. We observe that whenλ\\lambdais set to 0\.3, the model achieves the best performance in cross\-modal alignment\. More details are provided in Appendix[E](https://arxiv.org/html/2607.11978#A5)\.
### 4\.2\.Overall Performance under Joint Control
To the best of our knowledge, there is no unified molecular generation framework that jointly integrates gene expression profiles, structural design intentions, and chemical properties\. Therefore, we select baseline models based on their input modalities, including gene\-based, text\-based, and property\-guided approaches\. Since JoPMol operates under a multi\-condition setting, the comparison aims to evaluate the effectiveness of jointly modeling heterogeneous conditions rather than enforcing strict equivalence across inputs\.
Table[2](https://arxiv.org/html/2607.11978#S4.T2)presents a comprehensive comparison between JoPMol and multiple baseline models\. The results indicate that JoPMol outperforms all baselines on the six metrics related to structural consistency and sequence similarity\. Among the six metrics associated with chemical property optimization and overall statistics, JoPMol achieves the best performance on three metrics and the second\-best performance on the remaining three, demonstrating a clear overall advantage\. More specifically, compared with SOTA models such as TGM and GxVAEs, JoPMol improves the ECFP score by approximately 50%, indicating its ability to reliably preserve local chemical structures and design candidate molecules capable of inducing the target gene expression profiles\. For property optimization, even under multi\-condition control, JoPMol produces molecules whose property distributions closely match those of the test set using only numerical property values as conditioning signals, demonstrating effective property controllability and optimization\. Additionally, the overall statistics indicate that JoPMol achieves high validity and novelty and attains the best overall rank among all compared methods\. These results show that JoPMol exhibits stable design performance under jointly controlled settings\.
Furthermore, we validate the biological fidelity of the learned gene representations by reconstructing the latent features through the VAE decoder and comparing them with the original gene expression profiles\. The reconstructed profiles exhibit high consistency with the original ones in both distributional patterns and key expression characteristics, supporting the reliability and biological plausibility of the learned representations\. Detailed quantitative results and visualizations are reported in Appendix Figure[F\.6](https://arxiv.org/html/2607.11978#A6.F6)\.
We further evaluate the distributional quality of the designed molecules using RDKit and MACCS similarity metrics, as summarized in Appendix Table[F\.2](https://arxiv.org/html/2607.11978#A6.T2)\. The results demonstrate that JoPMol achieves the best overall performance, indicating close alignment with the empirical molecular distribution in terms of structural coverage and feature overlap\.
### 4\.3\.Transfer Tasks under Genetic Perturbations
Figure 2\.Representative examples of multi\-condition controlled molecular design under gene expression profiles, structural design intentions, and numerical property values for TP53 and CTSK\.To further evaluate the generalization capability and transfer stability of JoPMol, we extend JoPMol to ten ligand conditions and assess its performance accordingly\. Figure[2](https://arxiv.org/html/2607.11978#S4.F2)presents two representative cases corresponding to gene knockdown \(TP53\) and gene overexpression \(CTSK\) scenarios\. In each case, gene expression profiles induced by the target molecules, together with molecular design strategies, jointly guide the generation of candidate molecules, which are then compared against the target molecules\.
In the gene knockdown scenario, the designed molecules preserve key functional\-group patterns and achieve high structural similarity, with ECFP, APFP, and Dice scores of 0\.86, 0\.72, and 0\.93, indicating that JoPMol effectively captures molecular structural characteristics under unseen genetic perturbations\. Meanwhile, the three chemical properties of the candidate molecules remain highly consistent with those of the target molecules, demonstrating that JoPMol can effectively control chemical properties while preserving structural fidelity, enabling integrated molecular generation and optimization\. Additionally, in the gene overexpression scenario, JoPMol continues to design molecules with reasonable consistency in structural patterns and local functional\-group configurations, achieving ECFP, APFP, and Dice scores of 0\.68, 0\.77, and 0\.81, while maintaining strong property controllability\. Experimental results for the remaining eight transfer tasks are provided in Figure[F\.7](https://arxiv.org/html/2607.11978#A6.F7)\.
Overall, JoPMol maintains both structural fidelity and property optimization capability under different types of genetic perturbations, demonstrating its effectiveness on transfer tasks and its potential for precision molecular design\.
### 4\.4\.Ablation Studies on Joint Controllability
Table 3\.Ablation results at the overall statistics level under different fusion strategies and control settings\. W/concat, W/cross, W/FiLM, and W/adapter denote variants that replace the multimodal fusion in JoPMol with concatenation, cross\-attention, FiLM\-based modulation, and Adapter\-based fusion\. W/gene, W/struct, and W/prop correspond to gene\-controlled, structure\-controlled, and property\-controlled settings\.To analyze the effectiveness of different fusion strategies and the necessity of joint controllability, we conduct an extended ablation study including both fusion variants \(W/concat, W/cross, W/FiLM, and W/PropAdapter\) and single\-control settings \(W/gene, W/struct, and W/prop\)\. Table[3](https://arxiv.org/html/2607.11978#S4.T3)summarizes the results\.
Fusion Strategies\. We examine different fusion strategies and analyze their impact on joint controllability\. The concatenation strategy \(W/concat\) aggregates heterogeneous features without explicit interaction, achieving moderate performance \(Total: 0\.90, FCD: 16\.80\) but failing to capture cross\-modal dependencies\. Cross\-attention\-based fusion \(W/cross\) enables dynamic alignment between modalities and improves novelty to 94\.13, but leads to a substantially higher FCD of 27\.31, indicating degraded distributional consistency\. Feature\-wise linear modulation \(FiLM\)\(Perezet al\.,[2018](https://arxiv.org/html/2607.11978#bib.bib52)\)\(W/FiLM\) incorporates properties through feature\-wise transformations and achieves more balanced performance \(FCD: 16\.07\), suggesting improved stability\. W/adapter injects property conditions via a lightweight adapter module\(Houlsbyet al\.,[2019](https://arxiv.org/html/2607.11978#bib.bib53)\), enhancing novelty to 94\.24 but still resulting in a relatively high FCD of 23\.91, as it mainly performs local modulation without global guidance\. In contrast, JoPMol adopts a hierarchical fusion strategy by first jointly modeling biological and structural signals through bidirectional cross\-modal interaction, and subsequently incorporating property conditions as global guidance\. This enables stable integration of heterogeneous information and leads to improved performance in both structural consistency and distributional alignment\.
Control Settings\. We then analyze single\-control settings\. The results show that JoPMol exhibits advantages under joint condition control and achieves improved overall performance\. W/gene, which relies only on gene control, degrades noticeably in diversity and FCD\. W/struct, which uses only textual structural intentions, achieves the highest diversity \(90\.58\) but remains limited in Total and FCD\. W/prop, which uses only numerical chemical property values, attains the best novelty and a relatively low FCD \(12\.93\), but at the cost of reduced structural consistency and diversity\. In contrast, JoPMol achieves the best performance on total, FCD, and rank, indicating a more stable balance between design quality and distributional consistency\. Although its diversity is slightly lower than that of W/struct, this mainly results from the contraction of the feasible generation space under multiple control constraints\.
Figure 3\.Ablation results on joint controllability at different evaluation levels\. Left: structure consistency; Middle: sequence similarity; Right: property optimization\.Figure[3](https://arxiv.org/html/2607.11978#S4.F3)further presents a comparison between joint condition control and single control in terms of structure, sequence, and property\. For structural consistency, JoPMol achieves higher ECFP, APFP, and Dice scores than W/gene\. For sequence similarity, JoPMol shows improvements in BLEU and LCS while substantially reducing the Leven distance\. For property optimization, JoPMol performs comparably to W/prop on LogP, SA, and QED, with identical QED values, indicating strong alignment with the target distribution\. Overall, joint condition control preserves property optimization without sacrificing structural consistency and sequential similarity, enabling precision molecular design\.
Furthermore, we replace numerical chemical properties with textual property descriptions to investigate the impact of property representation on controllability\. The results show that using textual descriptions leads to inferior overall performance compared with numerical values\. Detailed results are provided in Table[F\.3](https://arxiv.org/html/2607.11978#A6.T3)\.
### 4\.5\.Case Study I: Property Controllability
Figure 4\.Evaluation of molecular distribution alignment and property controllability achieved by JoPMol\. \(a\) Distribution\-level comparison of logP, SA, and QED between the test molecules and the candidate molecules\. \(b\) Instance\-level analysis comparing a randomly selected target molecule with 1,000 designed candidates\.To further evaluate the chemical property controllability of JoPMol, Figure[4](https://arxiv.org/html/2607.11978#S4.F4)presents the distribution\-level and instance\-level results of the generated molecular properties\. Figure[4](https://arxiv.org/html/2607.11978#S4.F4)\(a\) compares the property distributions of the test\-set molecules \(gray curves\) with those of the generated molecules \(colored curves\)\. For all three properties, the generated distributions closely match the target distributions in both overall shape and mean values, indicating that the model can stably capture and reproduce property targets at the population level without noticeable distribution shift or mode collapse\. Additionally, the largest deviation is observed on SA, where the mean value shifts from 2\.70 \(target\) to 3\.12 \(generated\), corresponding to a relative deviation of approximately 15\.6%\. This deviation is reasonable given the relatively wide value range of SA, while the deviations of logP and QED remain much smaller\. These results indicate that JoPMol preserves the statistical structure of molecular property distributions while maintaining chemical plausibility\.
Figure[4](https://arxiv.org/html/2607.11978#S4.F4)\(b\) evaluates JoPMol at the instance level\. Given a randomly selected target molecule, 1,000 candidate molecules are designed under its property constraints\. The resulting property distributions are tightly concentrated around the target values, with mean values closely aligned with the specified conditions\. The relative deviations between the designed means and the target values for logP, SA, and QED are all controlled within approximately 4\-6%, demonstrating accurate and stable conditional generation\. Importantly, the narrow variance of the generated distributions suggests that JoPMol not only enforces property alignment but also reduces unnecessary structural dispersion\. This behavior reflects the effectiveness of the proposed multimodal fusion and conditional diffusion mechanisms in guiding molecular design\.
Overall, JoPMol achieves precise and reliable property controllability at both the distribution and instance levels, successfully balancing global distribution alignment with fine\-grained, target\-oriented optimization\. This capability is critical for precision molecular design and provides a strong foundation for customizable molecular generation in downstream drug discovery applications\.
### 4\.6\.Case Study II: Biological Simulation
Figure 5\.Docking visualization of the designed molecules under \(a\) breast cancer and \(b\) gastric cancer settings\. The predicted binding poses are shown within the binding pockets, with surrounding residues highlighted\. The designed molecules achieve binding affinities of \-8\.8kcal/mol\\mathrm\{kcal/mol\}on HER2 \(PDB ID: 3PP0\) and \-8\.6kcal/mol\\mathrm\{kcal/mol\}on tubulin \(PDB ID: 1JFF\)\.To further evaluate the practical applicability and biological relevance of JoPMol, we conduct molecular docking studies under both breast cancer and gastric cancer disease settings\.
Breast cancer\.We perform molecular docking against a representative breast cancer target, HER2 \(PDB ID: 3PP0\), to evaluate the biological plausibility of the designed molecules within a disease\-relevant context\. As illustrated in Figure[5](https://arxiv.org/html/2607.11978#S4.F5)\(a\), the designed molecule achieves a binding affinity of \-8\.8kcal/mol\\mathrm\{kcal/mol\}, outperforming a flavonoid\-like reference compoundCOC1=CC=C\(C=C1\)C2=C\(C\(=O\)C3=CC=CC=C3O2\)OC, which obtains \-8\.4kcal/mol\\mathrm\{kcal/mol\}under the same docking protocol\. The enlarged view further shows that the designed molecule is well accommodated within the binding pocket and exhibits favorable spatial complementarity with surrounding residues\. Specifically, the aromatic scaffold is positioned within the hydrophobic region of the pocket, enabling favorable van der Waals interactions, while the polar functional groups orient toward the pocket periphery, forming plausible hydrogen bonding interactions\. Importantly, no obvious steric clashes or unrealistic conformations are observed, indicating a chemically plausible binding mode\.
Gastric cancer\.To further assess the generalization ability of JoPMol beyond the training distribution, we conduct a case study in a gastric cancer setting\. Tubulin \(PDB ID: 1JFF\), a canonical target of taxanes widely used in gastric cancer therapy\(Maet al\.,[2023](https://arxiv.org/html/2607.11978#bib.bib8)\), is selected as the receptor\. Molecular docking is performed to evaluate the binding behavior of the designed molecules in a cross\-disease scenario\. As illustrated in Figure[5](https://arxiv.org/html/2607.11978#S4.F5)\(b\), the designed molecule occupies the binding pocket of the target protein and exhibits reasonable spatial complementarity with surrounding amino acid residues\. The docking result indicates a predicted binding energy of \-8\.6kcal/mol\\mathrm\{kcal/mol\}on the 1JFF target, suggesting favorable binding potential\. The enlarged view further visualizes the spatial interactions between the ligand and key residues, indicating a plausible binding conformation without apparent steric clashes or unstable orientations\.
These results collectively demonstrate that JoPMol not only produces biologically meaningful molecules consistent with the training domain, but also exhibits promising generalization capability to unseen disease contexts, highlighting its potential for cross\-disease precision molecular design\.
## 5\.Conclusion
In this study, we propose JoPMol, a jointly controlled molecular generative model for precision molecular design\. JoPMol integrates gene expression profiles, textual structural intentions, and property optimization within a unified diffusion\-based modeling framework\. By aligning multiple control signals, JoPMol enables controllable molecular generation and optimization\. Extensive experiments show that JoPMol achieves consistently improved performance compared with baseline methods across multiple controllability and quality metrics under realistic evaluation settings\.
Despite these results, JoPMol is currently evaluated primarily through in silico analyses and molecular docking, and the designed candidates have not yet been validated in wet\-lab experiments\. Future work will focus on extending experimental validation, expanding to broader disease targets and biological settings, and incorporating downstream biological screening to further examine the practical potential of jointly controlled precision molecular design in drug discovery\.
## References
- J\. S\. Angelo, I\. A\. Guedes, H\. J\. Barbosa, and L\. E\. Dardenne \(2023\)Multi\-and many\-objective optimization: present and future in de novo drug design\.Frontiers in Chemistry11,pp\. 1288626\.Cited by:[§2\.3](https://arxiv.org/html/2607.11978#S2.SS3.p1.1)\.
- P\. W\. Battaglia, J\. B\. Hamrick, V\. Bapst, A\. Sanchez\-Gonzalez, V\. Zambaldi, M\. Malinowski, A\. Tacchetti, D\. Raposo, A\. Santoro, R\. Faulkner,et al\.\(2018\)Relational inductive biases, deep learning, and graph networks\.arXiv preprint arXiv:1806\.01261\.Cited by:[§3\.1](https://arxiv.org/html/2607.11978#S3.SS1.p2.3)\.
- D\. Christofidellis, G\. Giannone, J\. Born, O\. Winther, T\. Laino, and M\. Manica \(2023\)Unifying molecular and textual representations via multi\-task language modelling\.InInternational Conference on Machine Learning,pp\. 6140–6157\.Cited by:[§2\.2](https://arxiv.org/html/2607.11978#S2.SS2.p1.1),[Table 2](https://arxiv.org/html/2607.11978#S4.T2.10.10.19.8.1)\.
- Y\. Deng, S\. S\. Ericksen, and A\. Gitter \(2025\)Chemical language model linker: blending text and molecules with modular adapters\.Journal of Chemical Information and Modeling65\(17\),pp\. 8944–8956\.Cited by:[§2\.2](https://arxiv.org/html/2607.11978#S2.SS2.p1.1)\.
- Q\. Duan, C\. Flynn, M\. Niepel, M\. Hafner, J\. L\. Muhlich, N\. F\. Fernandez, A\. D\. Rouillard, C\. M\. Tan, E\. Y\. Chen, T\. R\. Golub,et al\.\(2014\)LINCS Canvas Browser: interactive web app to query, browse and interrogate LINCS L1000 gene expression signatures\.Nucleic Acids Research42\(W1\),pp\. W449–W460\.Cited by:[Appendix A](https://arxiv.org/html/2607.11978#A1.p1.1),[1st item](https://arxiv.org/html/2607.11978#S4.I1.i1.p1.1)\.
- C\. Edwards, T\. Lai, K\. Ros, G\. Honke, K\. Cho, and H\. Ji \(2022\)Translation between molecules and natural language\.InEmpirical Methods in Natural Language Processing,pp\. 375–413\.Cited by:[§1](https://arxiv.org/html/2607.11978#S1.p1.1),[§1](https://arxiv.org/html/2607.11978#S1.p3.1),[§2\.2](https://arxiv.org/html/2607.11978#S2.SS2.p1.1),[Table 2](https://arxiv.org/html/2607.11978#S4.T2.10.10.18.7.1)\.
- H\. Gong, Q\. Liu, S\. Wu, and L\. Wang \(2024\)Text\-guided molecule generation with diffusion language model\.InAAAI Conference on Artificial Intelligence,pp\. 109–117\.Cited by:[§1](https://arxiv.org/html/2607.11978#S1.p1.1),[§2\.2](https://arxiv.org/html/2607.11978#S2.SS2.p1.1),[Table 2](https://arxiv.org/html/2607.11978#S4.T2.10.10.20.9.1)\.
- A\. Goyal and Y\. Bengio \(2022\)Inductive biases for deep learning of higher\-level cognition\.Proceedings of the Royal Society A478\(2266\),pp\. 20210068\.Cited by:[§3\.1](https://arxiv.org/html/2607.11978#S3.SS1.p2.3)\.
- S\. Guan and G\. Wang \(2024\)Drug discovery and development in the era of artificial intelligence: from machine learning to large language models\.Artificial Intelligence Chemistry2\(1\),pp\. 100070\.Cited by:[§2\.1](https://arxiv.org/html/2607.11978#S2.SS1.p1.1)\.
- J\. Hou \(2025\)De novo molecular design enabled by direct preference optimization and curriculum learning\.InJoint European Conference on Machine Learning and Knowledge Discovery in Databases,pp\. 95–111\.Cited by:[§2\.3](https://arxiv.org/html/2607.11978#S2.SS3.p1.1)\.
- N\. Houlsby, A\. Giurgiu, S\. Jastrzebski, B\. Morrone, Q\. De Laroussilhe, A\. Gesmundo, M\. Attariyan, and S\. Gelly \(2019\)Parameter\-efficient transfer learning for nlp\.InInternational conference on machine learning,pp\. 2790–2799\.Cited by:[§4\.4](https://arxiv.org/html/2607.11978#S4.SS4.p2.1)\.
- I\. Igashov, H\. Stärk, C\. Vignac, A\. Schneuing, V\. G\. Satorras, P\. Frossard, M\. Welling, M\. Bronstein, and B\. Correia \(2024\)Equivariant 3d\-conditional diffusion model for molecular linker design\.Nature Machine Intelligence6\(4\),pp\. 417–427\.Cited by:[§1](https://arxiv.org/html/2607.11978#S1.p3.1)\.
- T\. Inukai, A\. Yamato, M\. Akiyama, and Y\. Sakakibara \(2025\)Leveraging tree\-transformer vae with fragment tokenization for high\-performance large chemical model generation\.Communications Chemistry8\(1\),pp\. 228\.Cited by:[§1](https://arxiv.org/html/2607.11978#S1.p1.1),[§2\.3](https://arxiv.org/html/2607.11978#S2.SS3.p1.1),[Table 2](https://arxiv.org/html/2607.11978#S4.T2.10.10.17.6.1)\.
- W\. Jin, R\. Barzilay, and T\. Jaakkola \(2018\)Junction tree variational autoencoder for molecular graph generation\.InInternational conference on machine learning,pp\. 2323–2332\.Cited by:[§1](https://arxiv.org/html/2607.11978#S1.p3.1)\.
- K\. Kaitoh and Y\. Yamanishi \(2021\)TRIOMPHE: transcriptome\-based inference and generation of molecules with desired phenotypes by machine learning\.Journal of Chemical Information and Modeling61\(9\),pp\. 4303–4320\.Cited by:[§1](https://arxiv.org/html/2607.11978#S1.p1.1),[§1](https://arxiv.org/html/2607.11978#S1.p2.1),[§2\.1](https://arxiv.org/html/2607.11978#S2.SS1.p1.1),[Table 2](https://arxiv.org/html/2607.11978#S4.T2.10.10.12.1.1)\.
- S\. I\. Kang, J\. H\. Shin, B\. M\. Wu, and H\. S\. Choi \(2025\)Deep generative ai for multi\-arget therapeutic design: toward self\-improving drug discovery framework\.International Journal of Molecular Sciences26\(23\),pp\. 11443\.Cited by:[§1](https://arxiv.org/html/2607.11978#S1.p2.1)\.
- T\. Khater, S\. A\. Alkhatib, A\. AlShehhi, C\. Pitsalidis, A\. M\. Pappa, S\. T\. Ngo, V\. Chan, and V\. K\. Truong \(2025\)Generative artificial intelligence based models optimization towards molecule design enhancement\.Journal of Cheminformatics17\(1\),pp\. 116\.Cited by:[§1](https://arxiv.org/html/2607.11978#S1.p3.1)\.
- S\. Kim, P\. A\. Thiessen, E\. E\. Bolton, J\. Chen, G\. Fu, A\. Gindulyte, L\. Han, J\. He, S\. He, B\. A\. Shoemaker,et al\.\(2016\)PubChem substance and compound databases\.Nucleic Acids Research44\(D1\),pp\. D1202–D1213\.Cited by:[§3\.2](https://arxiv.org/html/2607.11978#S3.SS2.p1.5)\.
- D\. Kong, Y\. Huang, J\. Xie, E\. Honig, M\. Xu, S\. Xue, P\. Lin, S\. Zhou, S\. Zhong, N\. Zheng,et al\.\(2024\)Molecule design by latent prompt transformer\.Advances in Neural Information Processing Systems37,pp\. 89069–89097\.Cited by:[§1](https://arxiv.org/html/2607.11978#S1.p3.1)\.
- M\. Krenn, F\. Häse, A\. Nigam, P\. Friederich, and A\. Aspuru\-Guzik \(2020\)Self\-referencing embedded strings \(SELFIES\): a 100% robust molecular string representation\.Machine Learning: Science and Technology1\(4\),pp\. 045024\.Cited by:[§3\.2](https://arxiv.org/html/2607.11978#S3.SS2.p1.5)\.
- C\. Li and Y\. Yamanishi \(2025a\)AI\-driven transcriptome profile\-guided hit molecule generation\.Artificial Intelligence338,pp\. 104239\.Cited by:[§1](https://arxiv.org/html/2607.11978#S1.p1.1),[§1](https://arxiv.org/html/2607.11978#S1.p2.1),[§2\.1](https://arxiv.org/html/2607.11978#S2.SS1.p1.1),[§3\.3](https://arxiv.org/html/2607.11978#S3.SS3.p2.2),[Table 2](https://arxiv.org/html/2607.11978#S4.T2.10.10.14.3.1)\.
- C\. Li and Y\. Yamanishi \(2025b\)De novo generation of hit\-like molecules from gene expression profiles via deep learning\.InEuropean Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases,Cited by:[§2\.1](https://arxiv.org/html/2607.11978#S2.SS1.p1.1),[Table 2](https://arxiv.org/html/2607.11978#S4.T2.10.10.13.2.1)\.
- Q\. Liu, M\. Allamanis, M\. Brockschmidt, and A\. Gaunt \(2018\)Constrained graph variational autoencoders for molecule design\.Advances in Neural Information Processing Systems31\.Cited by:[§1](https://arxiv.org/html/2607.11978#S1.p3.1)\.
- H\. H\. Loeffler, J\. He, A\. Tibo, J\. P\. Janet, A\. Voronov, L\. H\. Mervin, and O\. Engkvist \(2024\)Reinvent 4: modern ai–driven generative molecule design\.Journal of Cheminformatics16\(1\),pp\. 20\.Cited by:[§1](https://arxiv.org/html/2607.11978#S1.p2.1)\.
- Y\. Luo, J\. Fang, S\. Li, Z\. Liu, J\. Wu, A\. Zhang, W\. Du, and X\. Wang \(2024\)Text\-guided small molecule generation via diffusion model\.iScience27\(11\)\.Cited by:[§1](https://arxiv.org/html/2607.11978#S1.p3.1)\.
- X\. Ma, Y\. Zhang, C\. Wang, and J\. Yu \(2023\)Efficacy and safety of combination chemotherapy regimens containing taxanes for first\-line treatment in advanced gastric cancer\.Clinical and Experimental Medicine23\(2\),pp\. 381–396\.Cited by:[§4\.6](https://arxiv.org/html/2607.11978#S4.SS6.p3.1)\.
- Y\. Matsukiyo, A\. Tengeiji, C\. Li, and Y\. Yamanishi \(2024\)Transcriptionally conditional recurrent neural network for de novo drug design\.Journal of Chemical Information and Modeling64\(15\),pp\. 5844–5852\.Cited by:[§1](https://arxiv.org/html/2607.11978#S1.p2.1),[§2\.1](https://arxiv.org/html/2607.11978#S2.SS1.p1.1)\.
- M\. Moret, L\. Friedrich, F\. Grisoni, D\. Merk, and G\. Schneider \(2020\)Generative molecular design in low data regimes\.Nature Machine Intelligence2\(3\),pp\. 171–180\.Cited by:[§1](https://arxiv.org/html/2607.11978#S1.p4.1)\.
- A\. Pathak, R\. Theagarajan, M\. M\. Rizqi, A\. S\. Nugraha, T\. Boruah, H\. Kumar, B\. Naik, S\. Yadav, A\. K\. Jha, A\. Trivedi,et al\.\(2025\)AI\-enabled drug and molecular discovery: computational methods, platforms, and translational horizons\.Discover Molecules2\(1\),pp\. 32\.Cited by:[§1](https://arxiv.org/html/2607.11978#S1.p1.1)\.
- Q\. Pei, W\. Zhang, J\. Zhu, K\. Wu, K\. Gao, L\. Wu, Y\. Xia, and R\. Yan \(2023\)BioT5: enriching cross\-modal integration in biology with chemical knowledge and natural language associations\.InEmpirical Methods in Natural Language Processing,pp\. 1102–1123\.Cited by:[§3\.2](https://arxiv.org/html/2607.11978#S3.SS2.p1.5)\.
- E\. Perez, F\. Strub, H\. De Vries, V\. Dumoulin, and A\. Courville \(2018\)Film: visual reasoning with a general conditioning layer\.InProceedings of the AAAI conference on artificial intelligence,Vol\.32\.Cited by:[§4\.4](https://arxiv.org/html/2607.11978#S4.SS4.p2.1)\.
- L\. Pinheiro Cinelli, M\. Araújo Marins, E\. A\. Barros da Silva, and S\. Lima Netto \(2021\)Variational autoencoder\.InVariational methods for machine learning with applications to deep networks,pp\. 111–149\.Cited by:[§2\.1](https://arxiv.org/html/2607.11978#S2.SS1.p1.1)\.
- D\. Shin, Y\. Son, D\. Lee, J\. Han, and T\. Kam \(2024\)Dynamic many\-objective molecular optimization: unfolding complexity with objective decomposition and progressive optimization\.InProceedings of the Thirty\-Third International Joint Conference on Artificial Intelligence,pp\. 6026–6034\.Cited by:[§2\.3](https://arxiv.org/html/2607.11978#S2.SS3.p1.1),[Table 2](https://arxiv.org/html/2607.11978#S4.T2.10.10.16.5.1)\.
- M\. Sun, J\. Xing, H\. Meng, H\. Wang, B\. Chen, and J\. Zhou \(2022\)Molsearch: search\-based multi\-objective molecular generation and property optimization\.InProceedings of the 28th ACM SIGKDD conference on knowledge discovery and data mining,pp\. 4724–4732\.Cited by:[§2\.3](https://arxiv.org/html/2607.11978#S2.SS3.p1.1),[Table 2](https://arxiv.org/html/2607.11978#S4.T2.10.10.15.4.1)\.
- \[35\]S\. Venkatraman, M\. Hasan, M\. Kim, L\. Scimeca, M\. Sendera, Y\. Bengio, G\. Berseth, and N\. MalkinOutsourced diffusion sampling: efficient posterior inference in latent spaces of generative models\.InForty\-second International Conference on Machine Learning,Cited by:[§3\.3](https://arxiv.org/html/2607.11978#S3.SS3.p2.2)\.
- S\. Wang, Y\. Guo, Y\. Wang, H\. Sun, and J\. Huang \(2019\)Smiles\-bert: large scale unsupervised pre\-training for molecular property prediction\.InProceedings of the 10th ACM international conference on bioinformatics, computational biology and health informatics,pp\. 429–436\.Cited by:[§1](https://arxiv.org/html/2607.11978#S1.p3.1)\.
- X\. Wang, C\. Wang, B\. Ji, J\. Wang, M\. Zheng, L\. Song, S\. Peng, and X\. Shang \(2026\)Multimodal pre\-training models of molecular representation for drug discovery\.National Science Review13\(1\),pp\. nwaf495\.Cited by:[§1](https://arxiv.org/html/2607.11978#S1.p4.1)\.
- Z\. Wang, C\. D\. Monteiro, K\. M\. Jagodnik, N\. F\. Fernandez, G\. W\. Gundersen, A\. D\. Rouillard, S\. L\. Jenkins, A\. S\. Feldmann, K\. S\. Hu, M\. G\. McDermott,et al\.\(2016\)Extraction and analysis of signatures from the gene expression omnibus by the crowd\.Nature Communications7\(1\),pp\. 12846\.Cited by:[3rd item](https://arxiv.org/html/2607.11978#S4.I1.i3.p1.1)\.
- W\. Weng, H\. Jiang, X\. Kong, and G\. Pau \(2025\)Text\-guided diverse\-expression diffusion model for molecule generation\.Chinese Physics B34\(5\),pp\. 050701\.Cited by:[§1](https://arxiv.org/html/2607.11978#S1.p3.1)\.
- M\. Yasunaga, J\. Leskovec, and P\. Liang \(2022\)LinkBERT: pretraining language models with document links\.InAssociation for Computational Linguistics,Cited by:[§3\.3](https://arxiv.org/html/2607.11978#S3.SS3.p3.5)\.
- X\. Zeng, F\. Wang, Y\. Luo, S\. Kang, J\. Tang, F\. C\. Lightstone, E\. F\. Fang, W\. Cornell, R\. Nussinov, and F\. Cheng \(2022\)Deep generative molecular design reshapes drug discovery\.Cell Reports Medicine3\(12\)\.Cited by:[§1](https://arxiv.org/html/2607.11978#S1.p1.1)\.
- Z\. Zhou, Y\. Li, P\. Hong, and H\. Xu \(2025\)Multimodal fusion with relational learning for molecular property prediction\.Communications Chemistry8\(1\),pp\. 200\.Cited by:[§1](https://arxiv.org/html/2607.11978#S1.p1.1)\.
Appendix
## Appendix ADataset Details
To support heterogeneous conditional molecule generation scenarios, we construct three transcriptomic datasets corresponding to different biological supervision signals, referred to as Data I, Data II, and Data III\. All datasets are uniformly preprocessed by retaining the 978 landmark genes defined in the LINCS L1000\(Duanet al\.,[2014](https://arxiv.org/html/2607.11978#bib.bib10)\)platform and applying feature\-wise normalization to ensure numerical comparability across conditions\. The raw sources of these datasets are publicly available under licenses permitting academic or research use\.
Data Iconsists of compound\-induced gene expression profiles derived from human cell lines under small\-molecule perturbations\. Each sample reflects a specific experimental setting characterized by the administered compound, dosage, exposure duration, and cellular context\. In this study, we focus on profiles collected from a single representative cell line, resulting in a large\-scale collection covering more than ten thousand unique compounds\. These expression signatures provide supervision signals that capture how molecular interventions reshape transcriptional states, enabling the model to associate chemical structures with downstream cellular responses\.
Data IIcontains target\-level perturbation signatures generated through genetic modulation, including knockdown or overexpression of individual protein targets\. Each profile encodes the transcriptomic consequence of manipulating a specific biological regulator\. We select a subset of biologically meaningful targets spanning kinase signaling, epigenetic regulation, and cancer\-related pathways\. These data enable conditioning molecule generation on desired target\-oriented functional effects, supporting controllable structure design driven by mechanistic intervention cues rather than purely chemical similarity\.
Data IIIis composed of disease\-associated transcriptomic signatures derived from patient\-level expression studies\. Each disease signature is obtained by aggregating multiple samples into a representative profile to reduce individual variability and enhance robustness\. In this work, only gastric cancer–related profiles are included\. This dataset allows the model to learn disease\-conditioned molecular generation, aligning chemical design with pathological transcriptional patterns\.
## Appendix BAnnotation Validation Protocol for Structural Design Intent
To ensure the reliability and consistency of the automatically extracted structural design intents, we adopt a multi\-stage human validation protocol involving independent expert review, consensus adjudication, and quality auditing\.
#### Automatic intent extraction\.
For each molecule, its SELFIES sequence is first processed by the BioT5 encoder to generate a candidate structural intent representation, which summarizes semantic patterns related to functional groups, substructures, and local chemical motifs\. These automatically generated intents serve as the initial annotations\.
#### Independent expert review\.
Two domain experts with formal training in medicinal chemistry independently examine each candidate intent without access to each other’s judgments\. Reviewers are provided with the original molecular structure \(rendered from SELFIES\), the predicted intent description, and a standardized checklist covering functional group correctness, structural completeness, chemical plausibility, and semantic consistency\. Each intent is labeled asaccepted,minor revision, orrejected, and optional correction notes are recorded when necessary\.
#### Consensus adjudication\.
All samples with conflicting labels or revision requests are jointly reviewed in a consensus meeting\. Reviewers discuss discrepancies, reconcile interpretations, and produce a finalized intent annotation\. If consensus cannot be reached, the sample is conservatively excluded from downstream training and evaluation to avoid introducing noisy supervision\.
#### Quality auditing and spot checking\.
A random subset of finalized annotations is periodically re\-evaluated to detect potential systematic bias or drift in annotation criteria\. This auditing step ensures long\-term consistency across the dataset and guards against latent labeling artifacts introduced during scaling\.
## Appendix CBaseline Models
To comprehensively evaluate JoPMol, we select representative baseline models from three complementary perspectives: \(i\) gene expression–conditioned molecular generation, \(ii\) text\-guided molecular generation, and \(iii\) chemically informed generative modeling\. All baselines are configured following their official implementations whenever possible and evaluated under consistent preprocessing and benchmarking settings\.
#### TRIOMRHE
TRIOMRHE is a Transformer\-based framework that learns the mapping between transcriptomic signatures and molecular responses through multi\-head attention and heterogeneous embedding strategies\. It models how gene expression patterns correspond to chemical structures at the representation level\. We follow the official configuration and apply it to our benchmark data\.
#### Gx2Mol\.
Gx2Mol is a hybrid neural architecture designed to generate molecules conditioned on transcriptomic signals\. It combines multiple neural components to encode gene expression features and translate them into molecular representations\. We apply the released implementation using our normalized gene profiles\.
#### GxVAEs\.
GxVAEs formulates transcriptome\-guided molecule generation under a conditional variational autoencoder framework, where gene expression vectors and molecular structures are coupled through a shared latent space\. This enables the model to reconstruct or sample molecules aligned with specific transcriptional states\. Pretrained weights are used for evaluation\.
#### MolSearch\.
MolSearch is a retrieval\-augmented molecular generation framework that formulates molecule discovery as a neural search problem over a learned chemical space\. It iteratively proposes candidate molecules and refines them based on similarity and property\-driven feedback, enabling efficient exploration of chemically valid regions\. We adopt the official implementation and use MolSearch as a representative search\-based baseline for property\-oriented molecular optimization\.
#### DyMol\.
DyMol is a dynamic molecular optimization framework that combines reinforcement learning with oracle\-guided evaluation to iteratively improve molecular candidates under target property constraints\. It employs adaptive policy updates to balance exploration and exploitation during optimization, enabling flexible control over multiple physicochemical objectives\. We use the standard configuration and include DyMol as a representative reinforcement learning\-based baseline for property\-driven molecular optimization\.
#### FRATTVAE
FRATTVAE is a variational autoencoder model that integrates recurrent encoders with attention mechanisms for molecular sequence modeling\. It learns latent molecular representations and reconstructs valid SMILES through attentive decoding, enabling controllable generation in latent space\. We adopt the standard implementation and use it as a representative chemically driven generative baseline\.
#### Mol\-T5\.
Mol\-T5 is a Transformer\-based large language model pretrained on large\-scale molecular SMILES corpora\. We employ the mol\-t5\-base checkpoint and prompt the model directly with target\-related textual descriptions to generate candidate molecules without additional fine\-tuning\.
#### Text\+ChemT5\.
This baseline integrates natural language prompts with chemically grounded representations using ChemT5, which is pretrained jointly on textual and molecular modalities\. Molecular descriptions and target cues are concatenated and provided as input, and generation is performed in a zero\-shot setting\.
#### TGM\-DLM
TGM\-DLM is a two\-stage diffusion framework for text\-conditioned molecular generation\. The first stage predicts a coarse molecular scaffold from textual input, while the second stage progressively refines it into a chemically valid molecule via conditional diffusion\. Official implementations are adopted for fair comparison\.
## Appendix DMetric Details
To provide a comprehensive evaluation of JoPMol, we report three groups of metrics: \(i\)*structure similarity*between generated and reference molecules \(ECFP, APFP, Dice\), \(ii\)*sequence similarity*between their SELFIES strings \(BLEU, Levenshtein, LCS\), and \(iii\)*overall statistics*that summarize generation quality at the set level \(Diversity, Total, and Rank\)\. These metrics jointly reflect structural fidelity, sequence\-level consistency, and overall generation reliability\.
#### ECFP Similarity \(↑\\uparrow\)\.
We compute the Tanimoto similarity between ECFP \(extended\-connectivity fingerprint\) of a generated moleculeggand its referencerr:
ECFP\(g,r\)=\|𝐟g∩𝐟r\|\|𝐟g∪𝐟r\|\.\\text\{ECFP\}\(g,r\)=\\frac\{\|\\mathbf\{f\}\_\{g\}\\cap\\mathbf\{f\}\_\{r\}\|\}\{\|\\mathbf\{f\}\_\{g\}\\cup\\mathbf\{f\}\_\{r\}\|\}\.Higher values indicate closer local substructure agreement\.
#### APFP Similarity \(↑\\uparrow\)\.
APFP \(atom\-pair fingerprint\) captures atom\-type pairs and their topological distances, emphasizing more global structural relations\. We compute the Tanimoto similarity in the same manner as ECFP, but on atom\-pair fingerprints\.
#### Dice Similarity \(↑\\uparrow\)\.
Dice similarity measures fingerprint overlap using:
Dice\(g,r\)=2\|𝐟g∩𝐟r\|\|𝐟g\|\+\|𝐟r\|,\\text\{Dice\}\(g,r\)=\\frac\{2\|\\mathbf\{f\}\_\{g\}\\cap\\mathbf\{f\}\_\{r\}\|\}\{\|\\mathbf\{f\}\_\{g\}\|\+\|\\mathbf\{f\}\_\{r\}\|\},which is more sensitive to shared features\.
#### BLEU \(↑\\uparrow\)\.
BLEU evaluatesnn\-gram overlap between the generated and reference SELFIES sequences, capturing local token\-level agreement\.
#### Levenshtein \(↓\\downarrow\)\.
Levenshtein distance counts the minimum number of edit operations required to transform the generated SELFIES into the reference sequence\. Lower values indicate better sequence consistency\.
#### LCS \(↑\\uparrow\)\.
LCS \(Longest Common Subsequence\) measures the length of the longest subsequence shared by two SELFIES strings while preserving order, reflecting long\-range sequential consistency\.
#### Diversity \(↑\\uparrow\)\.
Diversity quantifies the heterogeneity of the generated set by measuring average pairwise dissimilarity in fingerprint space \(higher is better\), indicating broader chemical coverage and reduced mode collapse\.
#### Total Score \(↑\\uparrow\)\.
Following our benchmark protocol, the overall Total score is defined as the product of three set\-level generation metrics:
Total=Validity×Novelty×Uniqueness\.\\text\{Total\}=\\text\{Validity\}\\times\\text\{Novelty\}\\times\\text\{Uniqueness\}\.This multiplicative form penalizes models that perform poorly in any single aspect and favors methods that simultaneously achieve reliable validity, sufficient novelty, and low redundancy\.
#### FCD \(↓\\downarrow\)\.
FCD \(Fréchet ChemNet Distance\) measures the distributional distance between generated molecules and reference molecules in the learned chemical feature space of a pretrained ChemNet model, reflecting the overall similarity of chemical properties and structural characteristics between the two molecule sets\.
#### Rank \(↓\\downarrow\)\.
To obtain a unified ranking across metrics, we first compute the per\-metric rank of each model \(with larger\-is\-better for↑\\uparrowmetrics and smaller\-is\-better for↓\\downarrowmetrics\)\. The final Rank score is then defined as the average of these ranks over the selected metric setℳ\\mathcal\{M\}:
Rank=1\|ℳ\|∑m∈ℳrankm,\\text\{Rank\}=\\frac\{1\}\{\|\\mathcal\{M\}\|\}\\sum\_\{m\\in\\mathcal\{M\}\}\\operatorname\{rank\}\_\{m\},whererankm\\operatorname\{rank\}\_\{m\}denotes the ranking position under metricmm\(ties, if any, are handled by assigning averaged ranks\)\. A lower Rank indicates better overall performance\.
## Appendix EImplementation Details\.
### E\.1\.Experiment Settings
Table E\.1\.Effect of the loss weighting coefficientλ\\lambdain JoPMol\. Moderate weighting \(λ=0\.3\\lambda=0\.3\) achieves the best trade\-off between structural similarity and distributional consistency\.All experiments are implemented in PyTorch with Python 3\.8 and trained on a workstation equipped with an NVIDIA A30 GPU\. Molecular structures are encoded as SELFIES sequences, with the maximum sequence length fixed at 256 tokens\.
Gene expression profiles are projected through a stack of three fully connected layers with hidden dimensions of 512, 256,128 and 64, respectively, yielding a compact latent representation for conditional control\. The same multilayer configuration is applied symmetrically during feature reconstruction\. The optimization process adopts a learning rate of1×10−41\\times 10^\{\-4\}and applies dropout with a ratio of 0\.2 to stabilize training\.
For structural semantic conditioning, tokenized SELFIES sequences are embedded into a 32\-dimensional trainable embedding space\. In parallel, molecular textual descriptions are encoded using a frozen language encoder with a 768\-dimensional hidden representation to extract high\-level semantic features, which are subsequently aligned with structural intent representations\.
The generative backbone follows a diffusion\-based formulation withT=2000T=2000denoising steps, optimized using a learning rate of1×10−41\\times 10^\{\-4\}and a dropout rate of 0\.1\. To further accelerate generation over remapped SELFIES sequences, a uniform skip\-sampling strategy is applied during the reverse diffusion process\. For the relevance alignment mechanism in JoPMol, the weighting coefficientλ\\lambdais fixed at 0\.3 and used consistently in all experiments\.
We further investigate the impact of the loss weighting coefficientλ\\lambda, which controls the contribution of the relevance\-guided objective in JoPMol\. As shown in Table[E\.1](https://arxiv.org/html/2607.11978#A5.T1), the choice ofλ\\lambdasignificantly influences the trade\-off between structural alignment, distributional consistency, and diversity\. Whenλ\\lambdais small \(e\.g\.,λ=0\.0\\lambda=0\.0or0\.10\.1\), the model relies primarily on the unconditional diffusion objective, resulting in high novelty and diversity but relatively weak structural alignment, as reflected by lower fingerprint similarities and BLEU scores\. Conversely, whenλ\\lambdabecomes large \(e\.g\.,λ≥0\.5\\lambda\\geq 0\.5\), the model is overly constrained by the relevance objective, which leads to degraded distributional consistency \(higher FCD\) and poorer sequence alignment \(higher Levenshtein distance\), despite maintaining high validity and novelty\. A moderate weighting \(λ=0\.3\\lambda=0\.3\) achieves the best balance across multiple criteria\. Specifically, it significantly improves structural similarity metrics \(Dice, APFP, and ECFP\) and sequence\-level alignment \(BLEU\), while also yielding the lowest FCD and Levenshtein distance\. These results indicate that an appropriate balance between conditional guidance and generative flexibility is crucial for achieving high\-quality and controllable molecular generation\.
### E\.2\.Use of Large Language Models
We use a large language model \(LLM\) solely for language polishing and grammar refinement to improve the clarity and readability of the manuscript\. The LLM is not involved in the generation of scientific content, experimental design, data analysis, model implementation, or result interpretation\. All technical contributions, experiments, and conclusions are conducted and verified by the authors\.
## Appendix FExperimental Supplements
### F\.1\.Additional Evaluation
Table F\.2\.Additional evaluation of JoPMol and baseline models on structural similarity, property consistency, and overall generation quality\.To further examine the robustness and practical utility of JoPMol, we conduct an additional evaluation using complementary metrics that explicitly measure structural similarity, property alignment, and overall generation quality\. As summarized in Table[F\.2](https://arxiv.org/html/2607.11978#A6.T2), the comparison covers representative transcriptome\-guided models \(TRIOMPHE, Gx2Mol, GxVAEs\), chemically driven generative models \(FRATTVAE\), text\-based generators \(Mol\-T5, Text\+ChemT5\), as well as a diffusion\-based baseline \(TGM\-DLM\)\.
#### Structural similarity\.
JoPMol achieves the strongest structural consistency across all evaluated fingerprint\-based metrics\. Specifically, JoPMol attains the highest RDKit Tanimoto similarity \(0\.39\) and MACCS similarity \(0\.61\), while simultaneously yielding the lowest Fréchet ChemNet Distance \(7\.40\), indicating superior alignment between the generated molecular distribution and the reference set\. Compared with the best competing baseline \(TGM\-DLM with FCD of 8\.94 and MACCS of 0\.54\), JoPMol reduces distributional discrepancy by approximately 17% and improves substructure overlap by more than 13%, demonstrating its ability to preserve chemically meaningful structural patterns under multi\-conditional guidance\.
#### Overall generation quality\.
JoPMol maintains perfect validity \(100\.0%\) and the highest uniqueness \(98\.91%\) among all compared approaches, confirming stable decoding and low redundancy during generation\. Although novelty is slightly lower than that of purely text\-driven models, it remains at a competitive level \(94\.29%\), indicating that JoPMol balances exploration and fidelity rather than over\-optimizing novelty at the cost of structural plausibility\. Notably, several baselines exhibit trade\-offs between validity and novelty, whereas JoPMol consistently preserves high\-quality generation across all three criteria\.
Overall, these results demonstrate that JoPMol achieves a favorable balance between structural fidelity, property alignment, and generative reliability\. By jointly integrating transcriptomic signals, semantic structural intent, and chemical property constraints, the model effectively avoids the common failure modes observed in single\-modality baselines, such as structural drift, property mismatch, or reduced validity\. The supplementary evaluation therefore corroborates the primary experimental conclusions and highlights the practical advantages of multi\-source controllable molecular generation\.
### F\.2\.Gene Expression Reconstruction via VAE\.
Figure F\.6\.Kernel density distributions of raw and reconstructed gene expression profiles for three randomly selected patients\. Reconstruction is obtained by encoding the profiles into the latent space and decoding them through the VAE, demonstrating close distributional alignment\.To verify that the learned latent representations preserve informative transcriptomic patterns, we conduct a reconstruction\-based validation using the gene expression encoder–decoder architecture\. Specifically, three patient\-level gene expression profiles are randomly sampled from the test set and projected into the latent space through the encoder\. The corresponding latent embeddings are then decoded back into reconstructed gene expression profiles\.
Figure[F\.6](https://arxiv.org/html/2607.11978#A6.F6)visualizes the empirical distributions of the original and reconstructed gene expression values for the three samples using kernel density estimation\. Across all cases, the reconstructed distributions closely match the raw distributions in terms of central tendency, spread, and overall shape, indicating that the encoder effectively captures the dominant statistical characteristics of high\-dimensional transcriptomic signals\.
Notably, the reconstructed profiles preserve both the unimodal structure and the tail behavior of the original distributions, suggesting that the latent bottleneck does not excessively smooth or distort biologically meaningful variability\. Minor deviations observed at extreme value ranges are expected due to stochastic sampling and regularization effects inherent to variational modeling\.
These results provide qualitative evidence that the learned latent representations maintain sufficient fidelity for downstream conditional molecular generation, supporting their use as reliable conditioning signals in JoPMol\.
Table F\.3\.Comparison between numerical chemical property values \(Value\) and textual property descriptions \(Text\) as conditional inputs under the same experimental setting\.Figure F\.7\.Qualitative comparison between known ligands and generated molecules for eight additional targets \(AKT1, AKT2, AURKB, EGFR, HDAC1, MTOR, PIK3CA, and SMAD3\)\.
### F\.3\.Chemical Property Representation Analysis
Table[F\.3](https://arxiv.org/html/2607.11978#A6.T3)compares the performance of using numerical chemical property values \(Value\) versus textual property descriptions \(Text\) as conditional inputs\. Both settings achieve similarly high validity and uniqueness, indicating comparable basic generation feasibility\. However, the Value setting consistently exhibits more stable performance on structure\- and distribution\-related metrics\. In particular, it achieves a lower FCD score than the Text setting \(11\.95 vs\. 12\.22\), suggesting closer alignment with the target distribution\. Comparable or slightly improved results are also observed on fingerprint similarity metrics, including MACCS, RDK, and ECFP\. In addition, the Value setting achieves a higher QED score \(0\.62 vs\. 0\.61\), reflecting more reliable preservation of drug\-likeness\.
In contrast, although the Text setting shows marginally better BLEU performance, this improvement does not translate into consistent gains in structural or property\-level quality and may introduce additional semantic ambiguity and noise\. Overall, directly conditioning on numerical chemical properties provides more precise, controllable, and stable optimization signals, and is therefore adopted in our model design\.
### F\.4\.Additional Cross\-Ligand Results
Figure[F\.7](https://arxiv.org/html/2607.11978#A6.F7)further reports qualitative generation results for the remaining eight ligand conditions, including AKT1, AKT2, AURKB, EGFR, HDAC1, MTOR, PIK3CA, and SMAD3\. For each target, the model receives transcriptomic perturbation profiles together with design intents as conditioning signals, and generates candidate molecules that are compared against the corresponding reference ligands in terms of structural similarity and physicochemical consistency\.
Across all eight cases, JoPMol consistently preserves core scaffold patterns and key functional groups, yielding moderate to high structural similarity measured by ECFP \(ranging approximately from 0\.51 to 0\.84\)\. Targets such as MTOR, PIK3CA, and SMAD3 exhibit particularly strong structural transferability, reflecting the model’s ability to capture conserved substructure motifs under distinct biological perturbations\. Even for more challenging targets with lower similarity scores \(e\.g\., AKT1 and AKT2\), the generated molecules still maintain chemically meaningful backbone alignment rather than degenerate or trivial structures\.
In addition to structural fidelity, the generated molecules remain well aligned with the reference ligands in major physicochemical properties\. The deviations in LogP, synthetic accessibility \(SA\), and drug\-likeness \(QED\) are generally small across all targets, indicating stable control over lipophilicity, synthetic feasibility, and overall molecular quality\. Notably, JoPMol avoids excessive property drift even when moderate structural variations are introduced, suggesting a balanced trade\-off between transfer flexibility and property preservation\.
Overall, these supplementary cases demonstrate that JoPMol maintains reliable cross\-ligand generalization and transfer stability beyond the examples reported in the main text\. The model consistently generates chemically plausible candidates that jointly satisfy biological conditioning, structural coherence, and property constraints, supporting its robustness for precision molecular design in diverse target settings\.Similar Articles
Controllable Molecular Generative Foundation Models
Proposes CoMole, a controllable molecular generative foundation model using motif-aware graph diffusion and reinforcement learning, achieving superior controllability across materials and drug discovery benchmarks.
Generating Developable 3D Molecules via Pocket-Conditioned Diffusion and Property-Aware Optimization
This paper introduces a novel diffusion-based generative model for structure-based drug design that decouples pocket and ligand representation learning and incorporates multi-scale interaction signals and property-aware optimization to generate developable 3D molecules with improved binding affinity and ADMET properties.
DrugGen 2: A disease-aware language model for enhancing drug discovery
DrugGen-2 fine-tunes GPT-2 using supervised learning and reinforcement learning (GRPO) to generate small molecules conditioned on both disease ontology and target protein sequences, achieving superior diversity and binding affinity for drug discovery.
ToolMol: Evolutionary Agentic Framework for Multi-objective Drug Discovery
ToolMol is an evolutionary agentic framework that combines a multi-objective genetic algorithm with an LLM-based operator to design small-molecule drugs, achieving state-of-the-art binding affinity and drug-likeness on multiple protein targets.
From Holo Pockets to Electron Density: GPT-style Drug Design with Density
This paper introduces EDMolGPT, an autoregressive framework that generates 3D molecular conformations from low-resolution electron density point clouds, improving structure-based drug design by leveraging physically meaningful density signals.