OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software
Summary
OSWorld-Science introduces a benchmark and evaluation environment for computer-using VLM agents on scientific software, covering 146 expert-designed tasks across domains like molecular drawing, pathology imaging, statistics, and physical simulation with artifact-based evaluators and a special agent harness.
View Cached Full Text
Cached at: 10/01/26, 08:23 AM
Paper page - OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software
Source: https://huggingface.co/papers/2609.39903 Published on Sep 30
·
Submitted byhttps://huggingface.co/iLOVE2D
Tianyuon Oct 1
Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Scientificsoftwarepresentsademandingtestforcomputer-usingagentsbasedonvisuallanguagemodels(VLMs):completingaresearchworkflowrequiresinterpretingspecializedinterfaces,manipulatingscientificobjects,andproducingverifiableresults.WethusintroduceOSWorld-Science,abenchmarkandevaluationenvironmentthatcombinesscientificallymeaningfultasks,artifact-basedevaluation,andanefficientagentharnessforstudyingcomputeruseinthescientificdomain.Thebenchmarkcontains12VLMsand146high-qualitytasksacrossseveralscientificdomainsandsoftwareconfigurations,coveringworkflowssuchasmoleculardrawingandretrosynthesis,pathologyimageanalysis,statisticalcomputing,andphysicalsimulation.Tasksaredevelopedthroughexpertproposalsanditerativehuman--AIco-design,withselectionguidedbyscientificvalueanddifficulty.Task-specificexecution-basedevaluatorsinspectapplicationstatesandgeneratedartifacts,includingmolecularstructures,segmentationmasks,plots,andnumericalresults,andawardpartialcreditforincompleteoutcomes.Ourspecialharnessintegratesmodeladapters,interaction-loopcontrol,andtrajectoryloggingtosupportcomparisonsofmodelsandinteractionstrategies.Ourresultsshowthatcurrentstate-of-the-artVLMswithastrongharnessstillfacechallengesinaddressingkeyquestionsinthescientificdomains.Wealsoanalyzethebenchmarkingresultsacrossmulti-linguistics,reasoningefforts,contextlengthandotherfactorsandderiveseveralimportantconclusionsanddirectionstoassistfuturedevelopment.Overall,weprovideanintegratedframeworkconnectingexpert-definedscientificgoalstoverifiablesoftwareoutcomes,enablingsystematicevaluationofbothagentcapabilitiesandharnessdesigninscientificworkflows.
View arXiv pageView PDFProject pageGitHub4Add to collection
Get this paper in your agent:
hf papers read 2609\.39903
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.39903 in a model README.md to link it from this page.
Datasets citing this paper1
#### SciAILab/OSWorld-Science-data Updatedabout 2 hours ago • 429 • 2
Spaces citing this paper1
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
OSWorld 2.0 is a new benchmark for evaluating computer-use agents on 108 long-horizon, real-world workflows. Current agents like Claude Opus 4.8 and GPT-5.5 achieve low completion rates, highlighting significant limitations in handling complex, multi-step tasks.
OpenComputer: Verifiable Software Worlds for Computer-Use Agents
OpenComputer presents a framework for creating verifiable software environments for computer-use agents, integrating state verifiers, self-improving verification layers, task synthesis, and evaluation systems across 33 desktop applications. Experiments show its verifiers align better with human judgment than LLM-as-judge, and frontier agents struggle with end-to-end completion.
DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?
This paper introduces DSAgentBench, the first benchmark for evaluating autonomous agents on complete, multi-tool data-science workflows in real computer environments. Results show that even the strongest agent (Claude-4.6-Sonnet) achieves only 56.70% task success, while open-source agents remain below 1%, revealing a substantial capability gap.
OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents
OR-Space is a benchmark for evaluating large language model agents in industrial operations research workflows, focusing on multi-stage task lifecycles and persistent workspaces beyond simple text generation.
SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool Acquisition
Presents SciToolAgent-Evo, an ontology-aware self-evolving LLM agent for open-world scientific tool acquisition, along with the OpenSciToolBench benchmark of 900 realistic tasks. The agent uses an evolving memory and LinUCB-based bandit gate to dynamically explore and acquire novel tools.