Domain-Adapted Molecular Language Models for Efficient Search of Make-on-Demand Libraries
Summary
This paper benchmarks pretrained molecular language models for virtual screening and demonstrates that explicit domain adaptation improves their performance and sample efficiency across drug discovery, materials, and catalysis libraries.
View Cached Full Text
Cached at: 08/19/26, 10:29 AM
# Domain-Adapted Molecular Language Models for Efficient Search of Make-on-Demand Libraries Source: [https://arxiv.org/abs/2608.17567](https://arxiv.org/abs/2608.17567) [View PDF](https://arxiv.org/pdf/2608.17567) > Abstract:Pretrained molecular language models are increasingly used as molecular encoders for learning structure\-property relationships\. However, their practical suitability for molecular discovery within and beyond their pretraining domain remains unclear\. Herein, we systematically benchmark four molecular language models across six virtual molecular libraries spanning drug discovery, organic materials, and catalysis\. Native molecular language model embeddings show substantial variation in discovery performance across libraries, whereas molecular fingerprints provide a consistently strong and robust baseline\. Consistent with a potential domain\-representation mismatch, we show that explicit domain adaptation substantially improves representation performance\. Fine\-tuning molecular language model encoders on structures from the target virtual library consistently improves sample efficiency, with several adapted encoders emerging as the top\-performing representations across the benchmark tasks\. These results show that molecular representation quality depends strongly on the target domain and that explicit adaptation can improve the practical utility of molecular foundation models\. More broadly, our findings establish domain\-adapted molecular representations as a promising strategy for sample\-efficient adaptive decision making in virtual screening and self\-driving laboratories\. ## Submission history From: Felix Strieth\-Kalthoff \[[view email](https://arxiv.org/show-email/45f349fc/2608.17567)\] **\[v1\]**Tue, 18 Aug 2026 09:27:56 UTC \(15,438 KB\)
Similar Articles
Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints
This paper benchmarks general-purpose LLMs against specialized diffusion models for generating binding molecules under 3D spatial constraints, finding that LLMs show promise despite currently lagging behind state-of-the-art approaches.
Rethinking Molecular OOD Generalization via Target-Aware Source Selection
This paper introduces SCOPE-Bench, a benchmark for evaluating molecular out-of-distribution generalization, and POMA, a framework using reinforcement learning to select source domains for domain adaptation, achieving significant error reductions on 3D molecular models.
Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search
The paper introduces Large Discovery Model (LDM), a recurrent architecture that couples generative models with Bayesian non-parametric surrogates to guide uncertainty-aware search in scientific domains like molecules and proteins, achieving significant performance gains over existing methods.
MolEmb: Multimodal Large Language Models Can Be Strong Molecular Embedding Models
The paper introduces MolEmb, a lightweight framework that adapts multimodal large language models for general molecular embedding, enabling context-aware representations and cross-modal retrieval, along with a diagnostic benchmark MolCAR.
Generating Developable 3D Molecules via Pocket-Conditioned Diffusion and Property-Aware Optimization
This paper introduces a novel diffusion-based generative model for structure-based drug design that decouples pocket and ligand representation learning and incorporates multi-scale interaction signals and property-aware optimization to generate developable 3D molecules with improved binding affinity and ADMET properties.