OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software

Hugging Face Daily Papers Papers

Summary

OSWorld-Science introduces a benchmark and evaluation environment for computer-using VLM agents on scientific software, covering 146 expert-designed tasks across domains like molecular drawing, pathology imaging, statistics, and physical simulation with artifact-based evaluators and a special agent harness.

Scientific software presents a demanding test for computer-using agents based on visual language models (VLMs): completing a research workflow requires interpreting specialized interfaces, manipulating scientific objects, and producing verifiable results. We thus introduce OSWorld-Science, a benchmark and evaluation environment that combines scientifically meaningful tasks, artifact-based evaluation, and an efficient agent harness for studying computer use in the scientific domain. The benchmark contains 12 VLMs and 146 high-quality tasks across several scientific domains and software configurations, covering workflows such as molecular drawing and retrosynthesis, pathology image analysis, statistical computing, and physical simulation. Tasks are developed through expert proposals and iterative human--AI co-design, with selection guided by scientific value and difficulty. Task-specific execution-based evaluators inspect application states and generated artifacts, including molecular structures, segmentation masks, plots, and numerical results, and award partial credit for incomplete outcomes. Our special harness integrates model adapters, interaction-loop control, and trajectory logging to support comparisons of models and interaction strategies. Our results show that current state-of-the-art VLMs with a strong harness still face challenges in addressing key questions in the scientific domains. We also analyze the benchmarking results across multi-linguistics, reasoning efforts, context length and other factors and derive several important conclusions and directions to assist future development. Overall, we provide an integrated framework connecting expert-defined scientific goals to verifiable software outcomes, enabling systematic evaluation of both agent capabilities and harness design in scientific workflows.
Original Article
View Cached Full Text

Cached at: 10/01/26, 08:23 AM

Paper page - OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software

Source: https://huggingface.co/papers/2609.39903 Published on Sep 30

·

Submitted byhttps://huggingface.co/iLOVE2D

Tianyuon Oct 1

Authors:

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

Abstract

Scientificsoftwarepresentsademandingtestforcomputer-usingagentsbasedonvisuallanguagemodels(VLMs):completingaresearchworkflowrequiresinterpretingspecializedinterfaces,manipulatingscientificobjects,andproducingverifiableresults.WethusintroduceOSWorld-Science,abenchmarkandevaluationenvironmentthatcombinesscientificallymeaningfultasks,artifact-basedevaluation,andanefficientagentharnessforstudyingcomputeruseinthescientificdomain.Thebenchmarkcontains12VLMsand146high-qualitytasksacrossseveralscientificdomainsandsoftwareconfigurations,coveringworkflowssuchasmoleculardrawingandretrosynthesis,pathologyimageanalysis,statisticalcomputing,andphysicalsimulation.Tasksaredevelopedthroughexpertproposalsanditerativehuman--AIco-design,withselectionguidedbyscientificvalueanddifficulty.Task-specificexecution-basedevaluatorsinspectapplicationstatesandgeneratedartifacts,includingmolecularstructures,segmentationmasks,plots,andnumericalresults,andawardpartialcreditforincompleteoutcomes.Ourspecialharnessintegratesmodeladapters,interaction-loopcontrol,andtrajectoryloggingtosupportcomparisonsofmodelsandinteractionstrategies.Ourresultsshowthatcurrentstate-of-the-artVLMswithastrongharnessstillfacechallengesinaddressingkeyquestionsinthescientificdomains.Wealsoanalyzethebenchmarkingresultsacrossmulti-linguistics,reasoningefforts,contextlengthandotherfactorsandderiveseveralimportantconclusionsanddirectionstoassistfuturedevelopment.Overall,weprovideanintegratedframeworkconnectingexpert-definedscientificgoalstoverifiablesoftwareoutcomes,enablingsystematicevaluationofbothagentcapabilitiesandharnessdesigninscientificworkflows.

View arXiv pageView PDFProject pageGitHub4Add to collection

Get this paper in your agent:

hf papers read 2609\.39903

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.39903 in a model README.md to link it from this page.

Datasets citing this paper1

#### SciAILab/OSWorld-Science-data Updatedabout 2 hours ago • 429 • 2

Spaces citing this paper1

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

OpenComputer: Verifiable Software Worlds for Computer-Use Agents

Hugging Face Daily Papers

OpenComputer presents a framework for creating verifiable software environments for computer-use agents, integrating state verifiers, self-improving verification layers, task synthesis, and evaluation systems across 33 desktop applications. Experiments show its verifiers align better with human judgment than LLM-as-judge, and frontier agents struggle with end-to-end completion.