An open benchmark for machine learning-based polymer property prediction
Summary
The paper introduces Polymer Benchmark 2026, an open dataset for benchmarking machine learning methods in polymer property prediction across diverse architectures and properties.
View Cached Full Text
Cached at: 09/24/26, 09:35 AM
# An open benchmark for machine learning-based polymer property prediction Source: [https://arxiv.org/abs/2609.27036](https://arxiv.org/abs/2609.27036) [View PDF](https://arxiv.org/pdf/2609.27036) > Abstract:Polymer property prediction lacks open, standardized benchmarks that enable rigorous comparison of machine\-learning methods, with existing resources covering only a narrow fraction of polymer architectures, such as homopolymers\. We introduce Polymer Benchmark 2026 \(PolyBench26\), an open dataset comprising nearly 250,000 polymer\-property datapoints across eight physical properties, including data from experimental measurements, density functional theory, and molecular dynamics\. The benchmark supports four evaluation tasks across homopolymers and alternating, random, and block copolymers: in\-distribution property prediction, dataset\-size scaling, repeat\-unit complexity, and transfer to held\-out polymer architectures\. We compare language model, graph\-based, and descriptor\-based approaches and find graph\-based models provide the lowest errors in property prediction, retain their advantage across the evaluated training\-set sizes, and remain robust to increasing repeat\-unit complexity\. PolyBench26 provides a reproducible foundation for developing models for the increasingly complex polymer design space\. The PolyBench26 benchmark is available open\-source at[this https URL](https://github.com/rlearsch/PolymerBenchmark2026)\. ## Submission history From: Robert Learsch \[[view email](https://arxiv.org/show-email/4458999e/2609.27036)\] **\[v1\]**Tue, 22 Sep 2026 20:32:53 UTC \(1,463 KB\)
Similar Articles
Procedural Pretraining for Molecular Property Prediction
The paper introduces a three-stage training pipeline using procedural pretraining to improve molecular property prediction, showing enhanced performance under data scarcity by learning inductive biases from abstract generated data.
PolyFusionAgent: A Multimodal Foundation Model and Autonomous AI Assistant for Polymer Property Prediction and Inverse Design
PolyFusionAgent is a framework that combines a multimodal polymer foundation model (PolyFusion) with a tool-augmented, literature-grounded design agent (PolyAgent) for polymer property prediction and inverse design, enabling evidence-linked discovery.
A Large-Scale Dataset and Benchmark: Do Protein-Ligand Models Learn Binding Sites or Just Binding Likelihood?
Introduces InteractBind, a large-scale dataset and benchmark for fine-grained evaluation of protein-ligand models, focusing on binding-site localization and non-covalent interaction prediction. Evaluates eight existing models and finds limited binding-site localization despite strong binary binding prediction.
Benchmarking data-driven material models on the classic Treloar dataset
This paper benchmarks popular machine learning-based frameworks for hyperelastic constitutive modeling using the classic Treloar dataset, comparing their performance, cost, and trade-offs to provide practical guidance for researchers.
PolyBridgeBench: Benchmarking Multimodal LLMs for Physics-Grounded Bridge Design
PolyBridgeBench is a new executable benchmark for evaluating multimodal LLMs on physics-grounded bridge design tasks, revealing gaps between deterministic validity and dynamic success in structure synthesis and repair.