SciHorizon-eLab: 面向科学具身智能体可扩展基准测试的智能体协议至任务编译器

arXiv cs.AI 论文

摘要

SciHorizon-eLab 是一个智能体协议至任务编译器,用于将科学协议编译为具身任务以进行可扩展基准测试,并引入了一个包含300个认证任务的基准测试。

arXiv:2609.30971v1 Announce Type: new Abstract: Embodied agents offer a promising route to automating scientific experimentation, yet their progress is constrained by the lack of reliable and systematic evaluation environments. Existing simulation-based laboratory benchmarks rely heavily on manual task engineering, making it challenging to systematically compile diverse scientific protocols into executable and verifiable embodied tasks at scale. To address this challenge, we introduce SciHorizon-eLab, an agentic protocol-to-task compiler that formulates scientific embodied task construction as a compilation problem. Given a natural-language protocol of scientific experiments, SciHorizon-eLab progressively compiles laboratory protocols into semantic-preserving embodied tasks through semantic grounding, executable task synthesis, and multi-stage simulation-based certification. The system generates semantically grounded environments, executable manipulation programs, and step-level success specifications, while enabling reproducible generation of expert demonstrations and execution traces. Using this pipeline, we further construct \BenchName, a ready-to-use benchmark comprising 300 certified tasks across diverse laboratory operations. It supports HIL task execution, reproducible expert-demonstration generation, and ordered step-level evaluation. Across representative tasks, the strongest policy attains an average success rate of only 49.7%, with further evaluations revealing pronounced weaknesses in human and embodied agent coordination. We publicly release the code, benchmark data, and evaluation toolkit at https://github.com/SciHorizon-elab/SciHorizon-elab.
查看原文
查看缓存全文

缓存时间: 2026/09/28 09:53

# 1 Introduction
Source: [https://arxiv.org/html/2609.30971](https://arxiv.org/html/2609.30971)
![[Uncaptioned image]](https://arxiv.org/html/2609.30971v1/figures/logo.png)

September 25, 2026

TECHNICAL REPORT

SciHorizon\-eLab: An Agentic Protocol\-to\-Task Compiler for Scalable Benchmarking of Scientific Embodied Agents

Maokai Qin1,2, Chuan Qin1,†,∗\{\}^\{1,\\dagger,^\{\*\}\}, Qi Zhang1,†, Dianyu Liu1, Zirui Liu1, Hongting Niu2, Yuanchun Zhou1, Hengshu Zhu1,∗\{\}^\{1,^\{\*\}\}

1Data Intelligence for Scientific Innovation Lab, Computer Network Information Center, Chinese Academy of Sciences

2Beihang University

†Project Co\-Lead∗Corresponding Authors

[qinmaokai@buaa\.edu\.cn](mailto:[email protected]);[chuanqin0426@gmail\.com](mailto:[email protected]);[zhangqi\.fqz@gmail\.com](mailto:[email protected]);[liudianyu00@gmail\.com](mailto:[email protected])

[202311998114@mail\.bnu\.edu\.cn](mailto:[email protected]);[niuhongting@buaa\.edu\.cn](mailto:[email protected]);[zyc@cnic\.cn](mailto:[email protected]);[zhuhengshu@gmail\.com](mailto:[email protected])

Keywords:Scientific embodied agents, protocol\-to\-task compilation, benchmarking, simulation

Contents

###### Abstract

Embodied agents offer a promising route to automating scientific experimentation, yet their progress is constrained by the lack of reliable and systematic evaluation environments\. Existing simulation\-based laboratory benchmarks rely heavily on manual task engineering, making it challenging to systematically compile diverse scientific protocols into executable and verifiable embodied tasks at scale\. To address this challenge, we introduceSciHorizon\-eLab, an agentic protocol\-to\-task compiler that formulates scientific embodied task construction as a compilation problem\. Given a natural\-language protocol of scientific experiments,SciHorizon\-eLabprogressively compiles laboratory protocols into semantic\-preserving embodied tasks through semantic grounding, executable task synthesis, and multi\-stage simulation\-based certification\. The system generates semantically grounded environments, executable manipulation programs, and step\-level success specifications, while enabling reproducible generation of expert demonstrations and execution traces\. Using this pipeline, we further construct SciVLABench, a ready\-to\-use benchmark comprising 300 certified tasks across diverse laboratory operations\. It supports human\-agent coordination task execution, reproducible expert\-demonstration generation, and ordered step\-level evaluation\. Across representative tasks, the strongest policy attains an average success rate of only 49\.7%, with further evaluations revealing pronounced weaknesses in human and embodied agent coordination\. We publicly release the code, benchmark data, and evaluation toolkit at[https://github\.com/SciHorizon\-elab/SciHorizon\-elab](https://github.com/SciHorizon-elab/SciHorizon-elab)\.

Recent advances in multimodal reasoning, foundation models, and robot learning have enabled embodied agents to interact with increasingly complex physical environments\[[1](https://arxiv.org/html/2609.30971#bib.bib16),[27](https://arxiv.org/html/2609.30971#bib.bib17),[6](https://arxiv.org/html/2609.30971#bib.bib18)\]\. Scientific laboratories are a particularly demanding domain for such agents: successful experimentation requires not only visuomotor manipulation, but also the interpretation of procedural instructions, an understanding of the functional roles of instruments and materials, the tracking of intermediate experimental states, and adherence to scientific and safety constraints over long action horizons\[[11](https://arxiv.org/html/2609.30971#bib.bib19),[22](https://arxiv.org/html/2609.30971#bib.bib20)\]\. As embodied agents emerge as promising assistants for scientific experimentation, their progress increasingly depends on reliable and systematic evaluation environments\.

Physics\-based simulation provides a safe, controllable, and scalable setting for evaluating scientific embodied agents\[[23](https://arxiv.org/html/2609.30971#bib.bib21),[19](https://arxiv.org/html/2609.30971#bib.bib22),[20](https://arxiv.org/html/2609.30971#bib.bib23)\], but constructing simulation\-ready laboratory tasks remains difficult\. Scientific protocols are written for trained human experimenters and typically describe procedures at a high level, whereas embodied benchmarks require explicit representations of entities, actions, states, constraints, and measurable success conditions\. Transforming a protocol involving operations such as transferring, heating, or mixing into an executable benchmark task therefore requires more than translating its instructions\. The protocol’s scientific intent must be grounded in a simulation environment, realized as an executable manipulation procedure, and expressed through objective, process\-level evaluation specifications\. Existing embodied and laboratory benchmarks commonly rely on manually authored or user\-configured task structures\[[13](https://arxiv.org/html/2609.30971#bib.bib5),[12](https://arxiv.org/html/2609.30971#bib.bib6),[15](https://arxiv.org/html/2609.30971#bib.bib7)\], while recent generative approaches automate individual components such as scene creation, task proposal, or controller synthesis\[[5](https://arxiv.org/html/2609.30971#bib.bib24),[2](https://arxiv.org/html/2609.30971#bib.bib14),[18](https://arxiv.org/html/2609.30971#bib.bib15)\]\. However, these components alone do not provide an end\-to\-end, protocol\-grounded construction process that jointly produces executable tasks and independent evaluation specifications, and then certifies their consistency through simulation before benchmark inclusion\.

To bridge this gap, we introduceSciHorizon\-eLab, an agentic protocol\-to\-task compiler that transforms natural\-language scientific protocols into executable and verifiable embodied tasks\. Its compilation pipeline comprises three stages: \(1\)Semantic\-Preserving Environment Compilation, which grounds a protocol in a simulator\-compatible environment and a structured task representation; \(2\)Executable Task Specification Generation, which synthesizes an executable manipulation program and an independently defined step\-level success specification; and \(3\)Simulation Certification and Task Instantiation, which validates structural correctness, physical executability, and ordered success conditions through simulation before instantiating certified tasks\. By unifying semantic grounding, executable synthesis, and simulation\-based certification,SciHorizon\-eLabenables scalable and reproducible construction of protocol\-grounded evaluation tasks\.

UsingSciHorizon\-eLab, we constructSciVLABench, a pre\-generated and ready\-to\-use benchmark for scientific embodied agents\. SciVLABench organizes 51 protocol\-grounded source scenarios into five primary operation families\. Using a library of 30 registered atomic manipulation skills, these scenarios are realized as 300 simulation\-certified tasks\. Each task contains a grounded environment, an executable reference program, and an ordered step\-level success specification, enabling reproducible, on\-demand generation of expert demonstrations and execution traces\. Certified randomization of poses, layouts, initial states, and visual conditions further yields diverse task instances without changing the task’s procedural or evaluation semantics\. In addition to autonomous execution tasks, SciVLABench includes 15 certified human\-agent collaborative tasks\. It supports binary task\-success assessment, ordered step\-level evaluation, and same\-task randomized evaluation\.

We conduct extensive experiments to evaluate both the benchmark and the underlying compiler\. On ten representative certified tasks, we compare three visuomotor policies using task\-level success and ordered step\-level completion; the strongest policy achieves an average task success rate of only 49\.7%, highlighting the difficulty of scientific manipulation\. We further assess human\-agent coordination tasks and transfer to geometrically distinct object bindings, exposing substantial limitations in current policies\. Finally, a controlled audit of 100 task commands measures the first\-pass reliability, recoverability, and remaining manual effort ofSciHorizon\-eLab\. To support reproducible research and rapid iteration, we publicly release the code, benchmark data, and evaluation toolkit at[https://github\.com/SciHorizon\-elab/SciHorizon\-elab](https://github.com/SciHorizon-elab/SciHorizon-elab)\. Our contributions are summarized as follows:

- ■\\blacksquareWe formulate scientific embodied task construction as a protocol\-to\-task compilation problem, providing a systematic path from natural\-language protocols to executable and objectively verifiable tasks\.
- ■\\blacksquareWe introduceSciHorizon\-eLab, an agentic compiler that integrates semantic\-preserving environment compilation, independently specified execution and evaluation logic, and multi\-stage simulation\-based certification\.
- ■\\blacksquareWe constructSciVLABench, a protocol\-grounded benchmark of 300 certified tasks, together with reproducible expert\-demonstration generation and evaluation protocols for procedural completion, transfer, and human\-agent coordination\.

Table 1:Comparison of representative embodied\-manipulation benchmarks and automated task\-construction systems\.Notes\.B/G denote benchmark/generator;//denote reported full, partial or related, and unreported support\. Scientific\-Laboratory \(Sci\. Lab\.\) denotes scientific\-laboratory specialization\. Ordered\-Step Evaluation \(Ord\. Step\) denotes ordered process\-level evaluation\. Binding Transfer \(Bind\. Transfer\) denotes evaluation with unseen object or instrument bindings under unchanged task semantics\. Human–Agent Coordination \(H–A Coord\.\) denotes coordination with hidden\-timing external interventions\. Protocol Input \(Protocol\) denotes traceable scientific\-protocol input\. Automatic Task Construction \(Auto Task\) denotes automatic construction of executable task definitions\. Independent Evaluation \(Indep\. Eval\.\) denotes success criteria generated separately from execution\. Simulation Certification \(Sim\. Cert\.\) denotes simulation validation before release\. Automatic Demonstration Generation \(Auto Demo\) denotes automatic generation of expert demonstrations\.

## 2Related Work

### 2\.1Scientific Embodied Agents

Automated laboratory systems are becoming an important pathway toward more reproducible and autonomous scientific experimentation\. Representative systems such as Chemputer\[[21](https://arxiv.org/html/2609.30971#bib.bib1)\], Artificial Chemist\[[7](https://arxiv.org/html/2609.30971#bib.bib2)\], Synbot\[[9](https://arxiv.org/html/2609.30971#bib.bib3)\], and MARS\-Chem\[[4](https://arxiv.org/html/2609.30971#bib.bib4)\]integrate robotic hardware with structured experimental workflows for chemical synthesis, materials exploration, and reaction optimization\. These systems demonstrate the feasibility of machine\-executed science, but are generally designed around specialized hardware, preconfigured procedures, and domain\-specific objectives\.

To reduce the cost and safety risks of physical experimentation, recent work has also explored scientific embodied agents in simulation\. Chemistry3D\[[13](https://arxiv.org/html/2609.30971#bib.bib5)\]supports robotic manipulation together with observable scientific states such as temperature, color, pH, transparent objects, and liquid interactions\. LabUtopia\[[12](https://arxiv.org/html/2609.30971#bib.bib6)\]provides interactive laboratory assets, procedural scene generation, and hierarchical evaluation from atomic manipulation to long\-horizon experiments\. Pipette\[[15](https://arxiv.org/html/2609.30971#bib.bib7)\]further offers editable wet\-lab assets, extensible task registration, and simulation\-based data augmentation\. These platforms provide valuable environments for scientific embodied learning, but their tasks are primarily manually specified or user\-configured rather than systematically compiled from natural\-language scientific protocols\.

### 2\.2Embodied Agent Benchmarking

Reliable benchmarking is essential for measuring the capabilities and failure modes of embodied agents\. Real\-robot systems such as RT\-1\[[1](https://arxiv.org/html/2609.30971#bib.bib16)\]and RT\-2\[[27](https://arxiv.org/html/2609.30971#bib.bib17)\]evaluate visuomotor policies in physical manipulation settings, but such evaluation is costly, difficult to standardize, and constrained by hardware and safety requirements\. Simulation therefore provides a more controllable and reproducible alternative\. RLBench\[[10](https://arxiv.org/html/2609.30971#bib.bib8)\]and LIBERO\[[14](https://arxiv.org/html/2609.30971#bib.bib9)\]provide standardized manipulation suites with expert demonstrations, ManiSkill2\[[8](https://arxiv.org/html/2609.30971#bib.bib10)\]studies generalization across object geometries and physical interactions, and CALVIN\[[16](https://arxiv.org/html/2609.30971#bib.bib11)\]evaluates long\-horizon skill composition\. VLABench\[[25](https://arxiv.org/html/2609.30971#bib.bib12)\]extends evaluation to visual, spatial, physical, and long\-horizon reasoning, while LabUtopia\[[12](https://arxiv.org/html/2609.30971#bib.bib6)\]focuses specifically on scientific laboratory environments\.

To reduce the cost of manually constructing tasks and demonstrations, RoboGen\[[24](https://arxiv.org/html/2609.30971#bib.bib13)\], RoboTwin 2\.0\[[2](https://arxiv.org/html/2609.30971#bib.bib14)\], and RoboGenesis\[[18](https://arxiv.org/html/2609.30971#bib.bib15)\]automate components such as task proposal, scene construction, execution\-program synthesis, and rollout collection\. However, as summarized in Table[1](https://arxiv.org/html/2609.30971#S1.T1), their task definitions and evaluation logic are not systematically derived from source scientific protocols, and successful execution alone does not guarantee semantic preservation, independently specified process\-level criteria, or consistent certification before benchmark inclusion\.SciHorizon\-eLabaddresses this gap by compiling scientific protocols into grounded environments, executable procedures, and independent step\-level success specifications, followed by simulation\-based certification\. Based on this framework, SciVLABench provides 300 certified scientific embodied tasks with reproducible expert demonstrations, ordered procedural evaluation, held\-out object\-binding evaluation, and human\-agent coordination tasks\.

![[Uncaptioned image]](https://arxiv.org/html/2609.30971v1/method_figure_2.png)

Figure 1:Overview of theSciHorizon\-eLabprotocol\-to\-task compilation framework\. Given a natural\-language scientific protocol,SciHorizon\-eLabprogressively compiles it into a simulator\-ready embodied task through semantic\-preserving environment compilation, executable task specification generation, and simulation\-based certification and task instantiation\.

## 3SciHorizon\-eLab

A scientific protocolppdescribes a laboratory procedure in natural language, including its experimental entities, ordered operations, and observable outcomes\. Our goal is to transform such a protocol into an embodied task that can be instantiated, executed, and automatically evaluated in a physics\-based environment\.

We formulate this protocol\-to\-task compilation problem as

𝖢⁡\(p\)=τ,\\mathsf\{C\}\(p\)=\\tau,\(1\)where𝖢\\mathsf\{C\}denotes the proposed compiler\. To support both execution and evaluation, the resulting taskτ\\taucontains three components:

τ=\(E,P,S\),\\tau=\(E,P,S\),\(2\)whereEErepresents the grounded simulation environment,PPspecifies the robot\-executable procedure, andSSspecifies the ordered conditions used to evaluate task completion\. The central challenge is to derive these mutually consistent components fromppwhile preserving its explicit scientific semantics and completing the robot\-specific details that are omitted from human\-oriented protocols\.

### 3\.1Framework Overview

As illustrated in Figure[1](https://arxiv.org/html/2609.30971#S2.F1),SciHorizon\-eLabtransforms a natural\-language scientific protocol into a simulator\-ready embodied task through three stages\. First,Semantic\-Preserving Environment Compilationinterprets the protocol and constructs the simulation environment required to realize its experimental procedure\. Second,Executable Task Specification Generationcompletes the robot\-specific operational details and independently derives the criteria used to evaluate task completion\. Third,Simulation Certification and Task Instantiationcompiles and executes the generated task in simulation, verifies its procedural outcomes, and instantiates qualified tasks under admissible randomization\. The following subsections describe these three stages in detail\.

### 3\.2Semantic\-Preserving Environment Compilation

The first stage converts the source protocolppinto an ordered sequence of grounded protocol stepsQQand a simulator\-compatible environmentEE\. We denote the output of this stage by

𝖦𝗋𝗈𝗎𝗇𝖽⁡\(p\)=\(Q,E\)\.\\mathsf\{Ground\}\(p\)=\(Q,E\)\.\(3\)The sequenceQ=\(q1,…,qm\)Q=\(q\_\{1\},\\ldots,q\_\{m\}\)preserves the procedural order and explicit experimental operations described inpp, whileEErecords the entities, asset bindings, relations, and initial\-state constraints required to realize these operations in simulation\. In this stage, semantic preservation requires that every task\-relevant element inQQandEEremains explicitly traceable to the source protocol\. The compiler retains the original entities, attributes, action descriptions, modifiers, spatial relations, and procedural constraints, while deferring robot\-specific execution details to the subsequent stage\. This separation ensures that low\-level manipulation decisions are not conflated with the scientific procedure specified by the protocol\.

A schema\-constrainedgrounding agentfirst identifies the entities and ordered operations explicitly described in the protocol\. Each grounded stepqiq\_\{i\}preserves its source operation, referenced entities, and stated constraints\. Mentions of the same experimental entity are assigned a consistent task\-level identifier across all steps, preventing unintended changes in object identity or functional role during protocol grounding\. The structured representation is accepted only if all entity references are successfully resolved and the original procedural order is strictly maintained\.

The grounded entities are subsequently instantiated into simulator representations according to their functional roles defined in the protocol\. Instruments, containers, and other manipulable objects are mapped to compatible simulator assets that support the required interactions\. Materials such as liquids and powders are modeled as contents or internal states associated with their physical carriers rather than as independent manipulable entities\. Asset retrieval and binding are constrained by each entity’s semantic category and interaction requirements\. The compiler may augment these representations with simulator\-specific properties, including collision geometry, interaction affordances, and initialization parameters, but such additions serve only to enable executable simulation and do not alter the protocol\-level entities, operations, or intended experimental outcomes\.

Finally, the grounded relations and initialization constraints are used to assemble the initial scene\. Deterministic checks ensure that all references remain consistent, selected assets support the required interactions, and initial spatial and state relations agree with the protocol\. The resulting pair\(Q,E\)\(Q,E\)therefore provides a traceable simulation representation of the source procedure\. The next stage uses this representation to determine how each grounded stepqiq\_\{i\}should be executed and how its completion should be evaluated\.

### 3\.3Executable Task Specification Generation

The second stage converts the grounded procedure and environment produced by the previous stage into the two task components required for execution and evaluation:

𝖲𝗉𝖾𝖼𝗂𝖿𝗒⁡\(Q,E\)=\(𝒫,𝒮\)\.\\mathsf\{Specify\}\(Q,E\)=\(\\mathcal\{P\},\\mathcal\{S\}\)\.\(4\)Let𝒫=\(P1,…,Pm\)\\mathcal\{P\}=\(P\_\{1\},\\ldots,P\_\{m\}\)and𝒮=\(s1,…,sm\)\\mathcal\{S\}=\(s\_\{1\},\\ldots,s\_\{m\}\), where𝒫\\mathcal\{P\}denotes the executable skill program and𝒮\\mathcal\{S\}denotes the step\-level success specification\. Each grounded stepqiq\_\{i\}serves as the shared alignment unit:Pi\{P\}\_\{i\}specifies how the step is executed, whilesis\_\{i\}defines the observable outcome required to verify its success\. This design ensures that execution and evaluation are consistently grounded in the same entities, operations, and procedural positions\.

#### 3\.3\.1 Executable Procedure Generation\.

Anexecution\-planning agentgenerates𝒫\\mathcal\{P\}using the complete grounded sequenceQQ, the entity and asset capabilities recorded inEE, and a registered library of atomic laboratory skills\. The agent plans over the full procedure rather than treating each step independently\. This allows it to track intermediate object and device states across steps and to insert robot\-specific operations that are implicit in the human\-oriented protocol, such as approaching, grasping, releasing, or repositioning an object\. These additions operationalize the stated procedure without changing its experimental step order or intended outcomes\.

EachPiP\_\{i\}is represented as an ordered sequence of parameterized atomic skills whose arguments refer to the shared task\-level entity identifiers\. A deterministic validator checks that the generated program preserves the order ofQQ, invokes only registered skills, supplies all required parameters, and uses entity bindings supported by the grounded environmentEE\.

#### 3\.3\.2 Success\-Condition Generation\.

In parallel, an independently promptedcondition\-planning agentderives𝒮\\mathcal\{S\}fromQQ,EE, and a registered library of verification predicates\. For eachqiq\_\{i\}, it converts the intended observable outcome into executable verification conditions over simulator states or events\. The resulting specifications capture multiple dimensions of task success, including spatial relations, object states, material or device states, and temporal constraints such as delayed execution, external triggers, or prevention of premature interactions\.

The condition\-planning agent does not have access to the generated program𝒫\\mathcal\{P\}\. Therefore, the verification conditions are derived solely from the grounded protocol rather than from the actions produced by the execution planner\. A deterministic validator verifies that each grounded stepqiq\_\{i\}is associated with a corresponding conditionsis\_\{i\}, that all predicate types and arguments are supported, and that all entity references remain consistent withEE\.

The two agents share the same grounded procedure and environment but independently generate their respective outputs\. This design preserves step\- and entity\-level alignment between𝒫\\mathcal\{P\}and𝒮\\mathcal\{S\}while decoupling action synthesis from success verification\. The resulting pair\(𝒫,𝒮\)\(\\mathcal\{P\},\\mathcal\{S\}\)is then passed to the certification stage, where simulation\-based execution and complementary verification procedures assess whether the generated tasks is executable, reproducible, and consistent with the source protocol\.

### 3\.4Simulation Certification and Task Instantiation

Given the taskτ\\tauand its aligned source information\(p,Q\)\(p,Q\), the final stage determines whether the generated specification forms a runnable and verifiable task in the MuJoCo physics engine\[[23](https://arxiv.org/html/2609.30971#bib.bib21)\]\. The stage first compiles and executes the task, then evaluates the resulting execution through complementary state\-based and visual\-semantic evidence\. Only tasks that pass all required checks are certified and used to generate benchmark instances\.

#### 3\.4\.1 Task Compilation and Physical Execution\.

A deterministic compiler transformsτ\\tauinto simulator\-native task and configuration classes\. Shared entity identifiers are preserved when binding assets, skill arguments, and verification conditions to simulator objects\. Before registration, the generated classes undergo deterministic validation to verify syntax correctness and required class and method interfaces\. Skill and condition parameters are produced by dedicated upstream planning stages and bound to simulator entities through preserved UID references\.

After successful registration, the program𝒫\\mathcal\{P\}is then executed inMuJoCo, producing an execution traceξ=\(x0,…,xH\)\\xi=\(x\_\{0\},\\ldots,x\_\{H\}\), wherextx\_\{t\}records the simulator state, skill\-execution status, robot waypoints, and multimodal observations at trace indextt\. Registration or execution failures prevent the current candidate from entering task verification and are recorded as structured diagnostics for subsequent inspection or refinement\.

#### 3\.4\.2 Complementary Task Verification\.

A successful execution is verified through two complementary verification channels\. First, the predicate verifier evaluates the specifications in𝒮\\mathcal\{S\}using simulator states recorded at the completion of each grounded stepqiq\_\{i\}\. State\-based predicates are checked at the corresponding step boundaries, whereas temporal predicates are evaluated over their associated execution segments\. The dependency relation≺\\precenforces that all predicates are satisfied in the procedural order required by the source protocol\. Second, avisual\-review agentassesses step\-aligned keyframes and multiview observations with respect to the source protocol and the expected outcomes of each step\. The keyframes include intermediate observations from the atomic skills associated with eachqiq\_\{i\}, as well as the final observation collected at the completion ofqiq\_\{i\}\.

Predicate verification checks properties explicitly represented in the simulator, whereas visual review assesses observable procedural semantics that the predicate library cannot fully capture\. Requiring both forms of evidence ensures that certified tasks are not only executable but also faithful to the intended scientific procedure\.

#### 3\.4\.3 Certification and Task Instantiation\.

Letδcomp\\delta\_\{\\mathrm\{comp\}\},δexec\\delta\_\{\\mathrm\{exec\}\},δpred\\delta\_\{\\mathrm\{pred\}\}, andδvis\\delta\_\{\\mathrm\{vis\}\}indicate successful compilation and registration, physical execution, predicate verification, and visual\-semantic review, respectively\. We define the certification result as

Γ⁡\(τ\)=𝕀⁡\[δcomp∧δexec∧δpred∧δvis\]\.\\Gamma\(\\tau\)=\\mathbb\{I\}\\left\[\\delta\_\{\\mathrm\{comp\}\}\\land\\delta\_\{\\mathrm\{exec\}\}\\land\\delta\_\{\\mathrm\{pred\}\}\\land\\delta\_\{\\mathrm\{vis\}\}\\right\]\.\(5\)A task is certified whenΓ⁡\(τ\)=1\\Gamma\(\\tau\)=1\.

Each certified taskτ\\taucan then generate task instances by sampling from its admissible initialization and randomization ranges\. The sampled parameters may vary object poses, workspace layouts, initial states, and visual conditions, while preserving the protocol\-level object roles, execution program, and success conditions\. The resulting certified tasks and randomized instances constitute the task resources described in Section[4](https://arxiv.org/html/2609.30971#S4)\.

## 4SciVLABench

UsingSciHorizon\-eLab, we construct SciVLABench, a ready\-to\-use benchmark for training and evaluating scientific embodied agents\. The benchmark consists of 300 certified tasks for autonomous execution and 15 human\-agent coordination tasks\.

Each certified task retains the grounded environment, executable expert program, and ordered success specifications generated bySciHorizon\-eLab\. Researchers can therefore directly instantiate tasks, generate expert demonstrations, and evaluate embodied policies without rerunning the complete protocol\-to\-task compilation process\. The remainder of this section describes the task sources and benchmark scope, followed by the released resources and supported evaluation protocols\.

### 4\.1Task Sources and Benchmark Scope

We curate laboratory procedures from scientific literature, automated laboratory systems, wet\-lab robotic benchmarks, and laboratory safety procedures\. We use a source scenario as the basic unit of task curation\. Each source scenario describes a coherent operation\-level procedure with explicitly identifiable entities, an ordered sequence of embodied actions, and observable intermediate or final outcomes\. A source scenario is retained when its entities and ordered operations can be explicitly identified, its core interactions can be grounded to available or extensible simulator assets and manipulation skills, and its outcomes can be evaluated through simulator states or visual observations\.

The resulting source pool contains 51 protocol\-grounded source scenarios into a taxonomy consisting of five primary operation families\. The five families include*Liquid Handling and Transfer*,*Mixing and Agitation*,*Solid Handling and Weighing*,*Thermal Control and Incubation*, and*Apparatus and Workspace Interaction*\. The operation families provide high\-level coverage of laboratory workflows\. This taxonomy is designed to characterize the operational scope of the benchmark rather than serve as an exhaustive categorization of all scientific experimentation\.

With a library of 30 registered atomic manipulation skills, the source scenarios are compiled into 300 simulation\-certified embodied tasks\. Each task can be instantiated under admissible variations of initial object poses, workspace configurations, and other task\-specific conditions while preserving its underlying operation semantics, executable procedure, and success specification\. This design enables repeated evaluation across diverse initial configurations without requiring task regeneration or manual redefinition\.

Human\-Agent Coordination Tasks\.Existing embodied laboratory benchmarks primarily focus on autonomous task execution, overlooking the coordination challenges that arise when experimental progress depends on external human interventions or environmental events\. To address this gap, our SciVLABench introduces a dedicated subset of 15 certified tasks for scalable evaluation of human\-agent coordination in scientific workflows\. These tasks are automatically compiled from protocol\-defined dependencies where execution cannot proceed until an externally triggered condition is satisfied, such as the ignition of a laboratory lamp or the addition of material to a container\. The agent must recognize the relevant event from its observations, safely defer interaction before the condition is met, and resume the remaining procedure only after the required state transition occurs\. This design enables controlled evaluation of coordination capabilities, including safe waiting, event understanding, intervention response, and reliable completion of interactive scientific procedures\.

### 4\.2Released Resources and Evaluation Protocols

Each certified task is released as a structured package containing a natural\-language instruction, a grounded environment specification, an executable expert program, ordered step\-level success specifications, and certification metadata\. By executing the expert program under admissible randomized initial conditions, the benchmark can reproducibly generate expert demonstrations, multimodal observations, execution traces, keyframes, and step\-level verification records\. This design enables on\-demand demonstration generation with different random seeds, avoiding dependence on a fixed set of pre\-recorded trajectories and supporting scalable evaluation under diverse task configurations\.

The ordered success specifications support evaluation at both the task and procedural levels\. Complete task success requires all step\-level conditions to be satisfied in the prescribed order, while step\-wise evaluation measures partial procedural completion when an agent fails before finishing the full task\. human\-agent coordination tasks additionally record premature interactions and post\-trigger responses, allowing pre\-intervention safety and event\-conditioned execution to be assessed separately from final task completion\. Formal metric definitions are provided in Section[5](https://arxiv.org/html/2609.30971#S5)\.

SciVLABench supports three complementary evaluation protocols\.*Same\-task randomized evaluation*trains and evaluates a policy on independently sampled instances of the same certified task, measuring robustness to variations in initial poses and workspace configurations\.*Held\-out object\-binding evaluation*replaces a training object with a geometrically distinct object while preserving the high\-level operation and success specification, measuring transfer across laboratory object bindings\.*human\-agent coordination evaluation*randomizes the timing of an external intervention and requires the policy to infer the resulting state change from its observations before resuming execution\.

## 5Experiments

We evaluate two complementary aspects of our work\. First, we assess the evaluation utility of SciVLABench by comparing representative visuomotor policies on certified laboratory tasks and examining their procedural completion, human\-agent coordination, and generalization to held\-out object bindings\. Second, we assess the practical reliability ofSciHorizon\-eLabthrough a controlled audit of its compilation and certification process\. We first describe the common experimental setup, followed by the benchmark evaluations and the construction audit\.

### 5\.1Experimental Setup

#### 5\.1\.1 Policies and Demonstrations\.

We evaluate three representative visuomotor policies:π0\.5\\pi\_\{0\.5\}\[[17](https://arxiv.org/html/2609.30971#bib.bib25)\], Action Chunking with Transformers \(ACT\)\[[26](https://arxiv.org/html/2609.30971#bib.bib26)\], and Diffusion Policy \(DP\)\[[3](https://arxiv.org/html/2609.30971#bib.bib27)\]\. For each task, we train a separate checkpoint using demonstrations generated by its certified expert program𝒫\\mathcal\{P\}\. Unless otherwise specified, each checkpoint is evaluated over 30 independently sampled episodes with randomized initial conditions\.

#### 5\.1\.2 Policy Inputs and Training\.

Each policy receives four RGB views—three fixed workspace cameras and one wrist\-mounted camera—together with robot proprioception\. All policies predict a 7D end\-effector command for position displacement, orientation displacement, and gripper control\. Images are resized to480×480480\\times 480forπ0\.5\\pi\_\{0\.5\}and256×256256\\times 256for ACT and DP\. ACT and DP operate at1010Hz, with ACT predicting chunks of 100 control steps\.

ACT and DP use AdamW for 100K optimization steps\. The learning rate is1×10−41\\times 10^\{\-4\}, and the batch size is 128\.π0\.5\\pi\_\{0\.5\}is fine\-tuned for 100K steps with a learning rate of2\.5×10−52\.5\\times 10^\{\-5\}and a global batch size of 64\. Privileged simulator states are used only to generate expert demonstrations and evaluate task outcomes; they are not provided as policy inputs\.

#### 5\.1\.3 Evaluation Metrics\.

We report task success rate \(SR\), which requires the complete ordered success specification𝒮\\mathcal\{S\}to be satisfied:

SR=1Neval∑n=1Neval𝕀\[Verify\(𝒮,ξ\(n\)∣≺\)=1\],\\mathrm\{SR\}=\\frac\{1\}\{N\_\{\\mathrm\{eval\}\}\}\\sum\_\{n=1\}^\{N\_\{\\mathrm\{eval\}\}\}\\mathbb\{I\}\\left\[\\operatorname\{Verify\}\\left\(\\mathcal\{S\},\\xi^\{\(n\)\}\\mid\\prec\\right\)=1\\right\],\(6\)whereNeval=30N\_\{\\mathrm\{eval\}\}=30,ξ\(n\)\\xi^\{\(n\)\}is the execution trace of episodenn, and≺\\precdenotes the required ordering among success conditions\.

To measure partial procedural completion, we additionally report step\-wise success rate \(SSR\):

SSR=1Neval​∑n=1Neval1m​∑i=1mzi\(n\),\\mathrm\{SSR\}=\\frac\{1\}\{N\_\{\\mathrm\{eval\}\}\}\\sum\_\{n=1\}^\{N\_\{\\mathrm\{eval\}\}\}\\frac\{1\}\{m\}\\sum\_\{i=1\}^\{m\}z\_\{i\}^\{\(n\)\},\(7\)wherezi\(n\)∈\{0,1\}z\_\{i\}^\{\(n\)\}\\in\\\{0,1\\\}indicates whether conditionsis\_\{i\}is satisfied in episodennafter all of its procedural prerequisites under≺\\prechave been satisfied, andmmis the number of ordered evaluation steps\. SR measures complete task execution, whereas SSR measures the fraction of the intended procedure completed before failure\. Both are shown as percentages\.

### 5\.2Benchmark Evaluation

We evaluate SciVLABench along three complementary dimensions: policy performance across heterogeneous laboratory procedures, coordination with externally triggered interventions, and generalization to held\-out object bindings\. Beyond ranking policies, these experiments examine whether the benchmark can expose procedural, coordination, and transfer failures that are obscured by aggregate task success alone\.

#### 5\.2\.1 Policy Comparison and Procedural Diagnosis

We evaluateπ0\.5\\pi\_\{0\.5\}, ACT, and Diffusion Policy on ten certified tasks spanning the five operation families defined in Section[4](https://arxiv.org/html/2609.30971#S4)\. The tasks range from basic object interaction to precision\-critical and multi\-stage laboratory manipulation\.

Table 2:Same\-task policy evaluation on ten tasks grouped by operation family\. Task names are shortened for space\.As shown in Table[2](https://arxiv.org/html/2609.30971#S5.T2), ACT achieves the highest average SR, whereasπ0\.5\\pi\_\{0\.5\}obtains the highest average SSR\. This difference indicates that the policy with the best complete\-task performance is not necessarily the one that makes the most procedural progress before failure\. Across the five operation families, performance does not exhibit a clear family\-level ordering and instead varies more strongly with task\-specific manipulation demands\. In particular, precision\-critical tasks such as*Unscrew The Pill Bottle*and*Insert Glass Rod Into Beaker*remain difficult for all three policies despite their short procedural\.

The gap between SR and SSR further reveals failures hidden by terminal success alone\. On*Transport and Mix*, complete success reaches at most 3\.3%, while SSR reaches 42\.5%, indicating that policies often complete early operations before failing at a later stage\. By combining family\-based coverage with ordered step\-level verification, SciVLABench supports both comparison across laboratory operation types and diagnosis of where execution breaks down within a procedure\.

#### 5\.2\.2 Human\-Agent Coordination

We further evaluate whether learned policies can coordinate robot behavior with external interventions whose timing is randomized and hidden from the policy\. We consider two representative scenarios\. In*Lamp\-Triggered Heating*, the robot must remain inactive until an alcohol lamp is ignited and then move the vessel to the heating region\. In*Post\-Addition Transport*, the robot must wait until material is added to a vessel before transporting it to the target workstation\. In both scenarios, the policy must identify the intervention from RGB observations alone without access to privileged state information\.

To distinguish pre\-intervention safety from post\-intervention behavior, we report task success rate \(SR\), premature violation rate \(PVR\), and post\-trigger response latency\. PVR measures the fraction of episodes in which the robot contacts or substantially moves a restricted object before the required intervention:

PVR=1Neval​∑n=1Nevalvn,vn∈\{0,1\}\.\\mathrm\{PVR\}=\\frac\{1\}\{N\_\{\\mathrm\{eval\}\}\}\\sum\_\{n=1\}^\{N\_\{\\mathrm\{eval\}\}\}v\_\{n\},\\quad v\_\{n\}\\in\\\{0,1\\\}\.\(8\)wherevnv\_\{n\}indicates a premature interaction in episodenn\.

Response latency measures the delay between the intervention trigger and the first valid robot response that initiates the subsequent manipulation stage:

Latency=tresponse∗−ttrigger,\\mathrm\{Latency\}=t^\{\*\}\_\{\\mathrm\{response\}\}\-t\_\{\\mathrm\{trigger\}\},\(9\)wheretresponse∗t\_\{\\mathrm\{response\}\}^\{\*\}denotes the timestamp of the first valid response\. The latency is averaged over episodes where a valid post\-trigger response is observed\.

Table 3:Human\-agent coordination under randomized and hidden intervention timing\. Latency is measured in seconds\. Each model–task pair is evaluated over 30 episodes\.Table[3](https://arxiv.org/html/2609.30971#S5.T3)reports results under randomized and hidden intervention timing\. Both policies achieve zero PVR across all evaluated episodes, demonstrating reliable avoidance of premature manipulation\. However, final task success remains below 17% for both tasks, with valid post\-trigger responses occurring after mean delays of 6\.8–8\.3 seconds\. This gap reveals that conservative waiting alone is insufficient for effective human–robot coordination: policies must not only defer actions safely but also detect interventions, respond appropriately, and complete the subsequent manipulation\. By disentangling pre\-trigger safety, post\-trigger responsiveness, and final task completion, SciVLABench reveals distinct coordination failures that cannot be captured by a single terminal success metric\.

#### 5\.2\.3 Generalization to Held\-Out Object Bindings

Same\-task evaluation varies initial conditions while retaining the same manipulated object\. We further examine whether policies can transfer to a geometrically distinct object binding that supports the same high\-level operation\. We consider two settings: lifting transfers from a beaker to a flask, while transport transfers from a cylinder to a Petri dish\. In each setting, the operation and success semantics remain unchanged, whereas object appearance, geometry, grasp affordances, and contact configurations differ\.

We quantify the object\-binding generalization gap as the performance drop from seen to held\-out bindings:

Δbind=SRseen−SRheld,\\Delta\_\{\\mathrm\{bind\}\}=\\mathrm\{SR\}\_\{\\mathrm\{seen\}\}\-\\mathrm\{SR\}\_\{\\mathrm\{held\}\},\(10\)where a smaller value indicates stronger retention\. Table[4](https://arxiv.org/html/2609.30971#S5.T4)reports the results on the seen and held\-out bindings\.

Table 4:Held\-out object\-binding generalization\. Each policy is trained only on the seen object and evaluated over 30 randomized episodes for each binding\. Results are SR \(%\); lowerΔbind\\Delta\_\{\\mathrm\{bind\}\}is better\.All three policies experience performance degradation when the manipulated object is replaced, indicating that robustness to randomized instances of seen objects does not necessarily translate into generalization to unseen object bindings\. The degradation is consistently larger for transport than for lifting, suggesting that object variation has a stronger impact when the policy must maintain stable grasping and control over a longer manipulation trajectory\.

Among the evaluated policies,π0\.5\\pi\_\{0\.5\}retains substantially more of its seen\-object performance, achieving an average held\-out SR of 76\.7%, compared with 45\.0% for ACT and 23\.3% for DP\. More broadly, the results demonstrate that the held\-out binding protocol isolates sensitivity to object geometry and affordances while preserving the underlying operation, complementing the initial\-state variations covered by same\-task evaluation\.

### 5\.3Compilation and Certification Audit

We conduct a controlled audit of 100 task commands to measure the first\-pass reliability, autonomous recoverability, and remaining manual effort ofSciHorizon\-eLab\. All commands use assets and operations covered by the registered libraries, allowing the audit to isolate task construction and certification within the current system scope rather than open\-world capability expansion\.

As shown in Table[5](https://arxiv.org/html/2609.30971#S5.T5), 58% of generated tasks pass structural, execution, predicate, and visual\-semantic checks on the first attempt\. For uncertified candidates, the certification workflow leverages structured diagnostics to trigger two automatic recovery mechanisms: simulator rerunning for transient execution failures and visual\-state correction for inconsistent initial configurations\. These mechanisms recover an additional 23% of candidates, raising the fully automated certification rate to 81%\.

This result demonstrates that certification inSciHorizon\-eLabis not merely a passive acceptance filter, but an active refinement process\. Through an agentic feedback loop, the system identifies recoverable execution and scene\-level inconsistencies and automatically improves generated tasks before benchmark admission\. In particular, visual\-state correction resolves mismatches between generated simulation scenes and textual task specifications, while predicate and visual\-semantic verification jointly ensure that corrected tasks preserve the intended procedural outcomes\. This autonomous recovery process substantially improves both task\-construction reliability and the scalability of benchmark generation\.

The remaining 19% of candidates fall outside the current automated recovery scope, primarily due to task\-specific spatial and interaction configurations not covered by the existing repair mechanisms\. These cases can be efficiently corrected through lightweight human intervention, requiring only 1–2 minutes on average per task\. This result demonstrates thatSciHorizon\-eLabenables scalable benchmark construction through reliable automated certification, with lightweight human intervention reserved only for rare task\-specific cases\.

Table 5:Controlled compilation and certification audit on 100 task candidates\. Recovery categories are mutually exclusive\. The cumulative rate shows the proportion certified after each recovery stage\.

## 6Conclusion

In this work, we introducedSciHorizon\-eLab, an agentic protocol\-to\-task compiler that formulated embodied task construction as a compilation problem\. By integrating semantic grounding, executable task synthesis, and multi\-stage simulation\-based certification,SciHorizon\-eLabtransformed natural\-language scientific protocols into semantic\-preserving, executable, and verifiable embodied tasks at scale\. Based on this framework, we developed SciVLABench, a certified benchmark with diverse laboratory scenarios that supports human\-in\-the\-loop interaction, reproducible expert demonstration generation, and fine\-grained step\-level evaluation\. Our experiments demonstrated that current embodied policies still face substantial challenges in scientific experimentation, particularly in long\-horizon manipulation and human\-agent coordination\. Beyond providing a benchmark,SciHorizon\-eLabestablishes a scalable paradigm for converting scientific knowledge into executable embodied experiences, paving the way for the development and evaluation of future scientific embodied agents\.

## Appendix ABenchmark Provenance and Operation\-Family Assignment

### A\.1Primary\-Family Assignment Principles

SciVLABenchcomprises 51 protocol\-grounded source scenarios curated from scientific literature, automated laboratory systems, wet\-lab robotic benchmarks, and laboratory\-safety procedures\. For benchmark organization, each scenario is assigned to one primary operation family based on its dominant embodied operation\.

The assignment is determined by the physical operation that primarily realizes the scientific objective, rather than by disciplinary context, laboratoryware type, or the secondary capabilities involved in the procedure\. For multi\-stage scenarios, the primary family is selected according to the operation that directly leads to the intended experimental outcome\. Other properties, including sensing requirements, safety\-critical execution, cultureware handling, long\-horizon composition, and human intervention, are retained as orthogonal metadata instead of being introduced as additional primary families\. The five primary operation families are defined as follows:

- ■\\blacksquareLiquid Handling and Transfer\.This family covers liquid handling operations, including the aspiration, dispensing, pouring, recovery, and transfer of liquid samples or reagents across containers, instruments, and testing media\. Sampling procedures involving pipettes, glass rods, pH paper, or colorimetric dishes are included when liquid transfer constitutes the dominant embodied operation\.
- ■\\blacksquareMixing and Agitation\.This family covers procedures whose primary objective is to homogenize, redistribute, or agitate materials through stirring, shaking, swirling, or repeated pouring\. A procedure is assigned to this family when achieving a mixed or homogenized state, rather than merely transferring material, constitutes the intended experimental outcome\.
- ■\\blacksquareSolid Handling and Weighing\.This family covers the access, manipulation, transfer, and quantitative handling of solid samples and their containers\. It includes operations such as weighing, powder transfer, material addition or removal for adjustment, and reagent\-container opening when access to the contents is required for subsequent solid\-sample handling\.
- ■\\blacksquareThermal Control and Incubation\.This family covers thermal\-control operations, including heating, incubation, drying, temperature monitoring, and post\-heating handling, whose primary purpose is to establish, maintain, monitor, or safely terminate a thermal condition\. These procedures may involve hot plates, magnetic heating devices, water baths, flames, drying ovens, thermometers, or insulating surfaces\.
- ■\\blacksquareApparatus and Workspace Interaction\.This family covers embodied operations involving the assembly, configuration, opening, closing, placement, storage, recovery, and organization of laboratory equipment, cultureware, tools, and workspace elements\. It also includes safety\-related recovery procedures when the dominant embodied operation concerns the state of the apparatus or workspace, rather than liquid handling, solid handling, mixing, or thermal processing\.

Table[S1](https://arxiv.org/html/2609.30971#A1.T1)summarizes the family\-level composition of the benchmark\. The complete scenario\-to\-family mapping is provided in Table[S2](https://arxiv.org/html/2609.30971#A1.T2)\.

Table S1:Family\-wise distribution of the 51 protocol\-grounded source scenarios in SciVLABench\.
### A\.2Scenario\-to\-Family Mapping

For compact presentation, we use the following abbreviations:LHTdenotes Liquid Handling and Transfer;MAdenotes Mixing and Agitation;SHWdenotes Solid Handling and Weighing;TCIdenotes Thermal Control and Incubation; andAWIdenotes Apparatus and Workspace Interaction\.

Table S2:Primary\-family assignments of the 51 protocol\-grounded source scenarios inSciHorizon\-eLab\.Scenario IDSource ScenarioPrimary FamilyS01Large\-Cylinder\-to\-Beaker Liquid TransferLHTS02Medium\-Volume Quantitative Liquid AdditionLHTS03Pipette\-to\-Tube DispensingLHTS04Sequential Dispensing Across Multiple TubesLHTS05Liquid Recovery from an Erlenmeyer Flask to a Graduated CylinderLHTS06Funnel\-Assisted Container TransferLHTS07Two\-Beaker Rinse TransferLHTS08Glass\-Rod Stirring in a BeakerMAS09Placement and Activation of a Magnetic StirrerMAS10Circular Mixing of an Erlenmeyer FlaskMAS11Tube\-Rack Retrieval and ShakingMAS12Repeated Pouring for Liquid MixingMAS13Gentle Swirling for Uniform Petri\-Dish CoatingMAS14Heating a Small Beaker on a Hot PlateTCIS15Heated Magnetic Stirring of an Erlenmeyer FlaskTCIS16Water\-Bath Incubation of a Test TubeTCIS17Flame Heating with Gentle Test\-Tube AgitationTCIS18Tripod\-Based Beaker Heating over an Alcohol LampTCIS19Drying a Petri Dish in a Drying OvenTCIS20Taring and Weighing a Petri DishSHWS21Pouring Solid Reagent from a Bottle into a Weighing DishSHWS22Sequential Weighing of Multiple PrecursorsSHWS23Transferring Weighed Powder into a BeakerSHWS24Unscrewing a Pill\-Bottle CapSHWS25Transferring Aspirin Samples from a Reagent BottleSHWS26Corrective Material Addition or Removal during WeighingSHWS27Assembly of a Gravity\-Filtration SetupAWIS28Precision Insertion of a Test Tube into a RackAWIS29Returning a Pipette to Its StandAWIS30Opening a Drying Oven and Loading a Petri DishAWIS31Retrieving and Storing a Reagent Bottle in a DrawerAWIS32Arranging a Tripod and Heat SourceAWIS33Pipette Sampling into a Petri DishLHTS34Opening and Closing a Petri\-Dish LidAWIS35Petri\-Dish Spotting and Lid ReplacementLHTS36Transferring a Test\-Tube Sample to a Petri DishLHTS37Transporting a Petri Dish to an Observation or Processing AreaAWIS38Pipette Spotting Followed by Tool StorageLHTS39Organizing a Test\-Tube ArrayAWIS40Measuring Beaker Temperature with a ThermometerTCIS41Sampling onto pH Paper with a PipetteLHTS42Transferring a Sample to pH Paper with a Glass RodLHTS43Mounting pH Paper and Dispensing a SampleLHTS44Dispensing a Test\-Tube Sample into a Colorimetric DishLHTS45Positioning and Returning a Contact pH ProbeAWIS46Wiping a Chemical Spill from the WorkbenchAWIS47Righting a Fallen Erlenmeyer FlaskAWIS48Moving a Hot Vessel onto an Insulating MatTCIS49Safely Removing a Test Tube after Flame HeatingTCIS50Closing an Equipment Door and DrawerAWIS51Slow Collision\-Aware Transport of GlasswareAWIAnalytical, sensing, cultureware\-handling, and safety\-critical properties exhibited by some scenarios are preserved as secondary task metadata rather than used to define primary operation families\. For example, thermometer\-based temperature measurement is assigned to TCI while retaining a sensing attribute; pH\-paper sampling is assigned to LHT while retaining an analytical\-testing attribute; and spill cleanup or apparatus recovery is assigned to AWI while retaining a safety\-critical attribute\. Human intervention is similarly modeled as an orthogonal coordination attribute rather than an additional primary operation family\.

## Appendix BRegistered Atomic Skill Library

SciHorizon\-eLableverages a registered library of 30 atomic manipulation skills to compile grounded protocol steps into robot\-executable programs\. Each skill defines a reusable simulator\-level operation parameterized by grounded entities, target poses, spatial offsets, or task\-specific execution settings\. The skill\-planning agent can compose multiple atomic skills to realize a single protocol step, while a deterministic validator ensures that each generated program invokes only registered skills with complete and valid arguments\.

Table[S3](https://arxiv.org/html/2609.30971#A2.T3)lists the complete atomic skill set used in the current implementation\. The skill names correspond to callable interfaces exposed by the simulator execution API\.

Table S3:Registered atomic manipulation skills used by SciHorizon\-Lab\. Skill names are reported using the exact callable identifiers in the simulator execution API\.IDSkillFunctional Description1step\_trajectoryExecutes a general robot trajectory represented as an ordered sequence of target waypoints\.2moveto\_entityMoves the robot end effector to a task\-defined position above a specified grounded entity\.3movetoMoves the robot end effector to a specified target position\.4pickGrasps and lifts a specified manipulable object\.5placePrecisely places a manipulated object inside a target container or receptacle\.6dropReleases a manipulated object onto the workspace surface\.7open\_doorOpens an articulated door\.8close\_doorCloses an articulated door\.9open\_drawerOpens an articulated drawer\.10pressPerforms a pressing action, such as activating a button or device control\.11pullApplies a pulling motion to a specified object or articulated component\.12pushApplies a pushing motion to a specified object or articulated component\.13pourPerforms a basic pouring motion by tilting a manipulated container\.14pour\_to\_entityMoves a source container above a specified target entity and performs a pouring operation\.15liftRaises the robot end effector or a currently manipulated object\.16resetReturns the robot to its predefined default configuration\.17close\_gripperCloses the robot gripper\.18open\_gripperOpens the robot gripper\.19flipFlips or turns over a manipulated object\.20rotateRotates the robot wrist or the currently manipulated object\.21open\_laptopOpens the articulated display of a laptop computer\.22waitKeeps the robot inactive for a specified number of simulation steps\.23wait\_forKeeps the robot inactive until a specified trigger time or task\-defined waiting condition is reached\.24shakeApplies a repeated shaking motion to a manipulated object\.25move\_offsetMoves the robot end effector by a specified positional offset relative to its current pose\.26insert\_to\_entityInserts a manipulated object into the opening or designated insertion region of a target entity\.27stir\_entity\_with\_toolUses a manipulated stirring tool to agitate the contents of a specified container\.28unscrew\_capUnscrews and removes the cap of a specified container\.29aspirateAspirates a liquid sample using a pipette, dropper, or compatible liquid\-handling instrument\.30dispenseDispenses or spots an aspirated liquid sample from a pipette, dropper, or compatible liquid\-handling instrument\.
## References

- \[1\]A\. Brohan, N\. Brown, J\. Carbajal, Y\. Chebotar, J\. Dabis, C\. Finn, K\. Gopalakrishnan, K\. Hausman, A\. Herzog, J\. Hsu,et al\.\(2023\)RT\-1: robotics transformer for real\-world control at scale\.InRobotics: Science and Systems XIX,External Links:[Document](https://dx.doi.org/10.15607/RSS.2023.XIX.025)Cited by:[§1](https://arxiv.org/html/2609.30971#S1.p1.1),[§2\.2](https://arxiv.org/html/2609.30971#S2.SS2.p1.1)\.
- \[2\]T\. Chen, Z\. Chen, B\. Chen, Z\. Cai, Y\. Liu, Z\. Li, Q\. Liang, X\. Lin, Y\. Ge, Z\. Gu, W\. Deng, Y\. Guo, T\. Nian, X\. Xie, Q\. Chen, K\. Su, T\. Xu, G\. Liu, M\. Hu, H\. Gao, K\. Wang, Z\. Liang, Y\. Qin, X\. Yang, P\. Luo, and Y\. Mu\(2025\)RoboTwin 2\.0: a scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation\.External Links:2506\.18088Cited by:[§1](https://arxiv.org/html/2609.30971#S1.p2.1),[§1](https://arxiv.org/html/2609.30971#S1.p6.1.9.1.1),[§2\.2](https://arxiv.org/html/2609.30971#S2.SS2.p2.1)\.
- \[3\]\(2023\)Diffusion policy: visuomotor policy learning via action diffusion\.InProceedings of Robotics: Science and Systems,Daegu, Republic of Korea\.External Links:[Document](https://dx.doi.org/10.15607/RSS.2023.XIX.026)Cited by:[§5\.1\.1](https://arxiv.org/html/2609.30971#S5.SS1.SSS1.p1.1)\.
- \[4\]T\. Dai, S\. Vijayakrishnan, F\. T\. Szczypiński, J\. Ayme, E\. Simaei, T\. Fellowes, R\. Clowes, L\. Kotopanov, C\. E\. Shields, Z\. Zhou, J\. W\. Ward, and A\. I\. Cooper\(2024\)Autonomous mobile robots for exploratory synthetic chemistry\.Nature635\(8040\),pp\.890–897\.External Links:[Document](https://dx.doi.org/10.1038/s41586-024-08173-7)Cited by:[§2\.1](https://arxiv.org/html/2609.30971#S2.SS1.p1.1)\.
- \[5\]M\. Deitke, E\. VanderBilt, A\. Herrasti, L\. Weihs, J\. Salvador, K\. Ehsani, W\. Han, E\. Kolve, A\. Farhadi, A\. Kembhavi, and R\. Mottaghi\(2022\)ProcTHOR: large\-scale embodied AI using procedural generation\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\.5982–5994\.Cited by:[§1](https://arxiv.org/html/2609.30971#S1.p2.1)\.
- \[6\]D\. Driess, F\. Xia, M\. S\. M\. Sajjadi, C\. Lynch, A\. Chowdhery, B\. Ichter, A\. Wahid, J\. Tompson, Q\. Vuong, T\. Yu, W\. Huang, Y\. Chebotar, P\. Sermanet, D\. Duckworth, S\. Levine, V\. Vanhoucke, K\. Hausman, M\. Toussaint, K\. Greff, A\. Zeng, I\. Mordatch, and P\. Florence\(2023\)PaLM\-E: an embodied multimodal language model\.InProceedings of the 40th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.202,pp\.8469–8488\.Cited by:[§1](https://arxiv.org/html/2609.30971#S1.p1.1)\.
- \[7\]R\. W\. Epps, M\. S\. Bowen, A\. A\. Volk, K\. Abdel\-Latif, S\. Han, K\. G\. Reyes, A\. Amassian, and M\. Abolhasani\(2020\)Artificial chemist: an autonomous quantum dot synthesis bot\.Advanced Materials32\(30\),pp\.2001626\.External Links:[Document](https://dx.doi.org/10.1002/adma.202001626)Cited by:[§2\.1](https://arxiv.org/html/2609.30971#S2.SS1.p1.1)\.
- \[8\]J\. Gu, F\. Xiang, X\. Li, Z\. Ling, X\. Liu, T\. Mu, Y\. Tang, S\. Tao, X\. Wei, Y\. Yao, X\. Yuan, P\. Xie, Z\. Huang, R\. Chen, and H\. Su\(2023\)ManiSkill2: a unified benchmark for generalizable manipulation skills\.InInternational Conference on Learning Representations,Cited by:[§2\.2](https://arxiv.org/html/2609.30971#S2.SS2.p1.1)\.
- \[9\]T\. Ha, D\. Lee, Y\. Kwon, M\. S\. Park, S\. Lee, J\. Jang, B\. Choi, H\. Jeon, J\. Kim, H\. Choi, H\. Seo, W\. Choi, W\. Hong, Y\. J\. Park, J\. Jang, J\. Cho, B\. Kim, H\. Kwon, G\. Kim, W\. S\. Oh, J\. W\. Kim, J\. Choi, M\. Min, A\. Jeon, Y\. Jung, E\. Kim, H\. Lee, and Y\. Choi\(2023\)AI\-driven robotic chemist for autonomous synthesis of organic molecules\.Science Advances9\(44\),pp\.eadj0461\.External Links:[Document](https://dx.doi.org/10.1126/sciadv.adj0461)Cited by:[§2\.1](https://arxiv.org/html/2609.30971#S2.SS1.p1.1)\.
- \[10\]S\. James, Z\. Ma, D\. R\. Arrojo, and A\. J\. Davison\(2020\)RLBench: the robot learning benchmark and learning environment\.IEEE Robotics and Automation Letters5\(2\),pp\.3019–3026\.External Links:[Document](https://dx.doi.org/10.1109/LRA.2020.2974707)Cited by:[§1](https://arxiv.org/html/2609.30971#S1.p6.1.3.1.1),[§2\.2](https://arxiv.org/html/2609.30971#S2.SS2.p1.1)\.
- \[11\]R\. D\. King, J\. Rowland, S\. G\. Oliver, M\. Young, W\. Aubrey, E\. Byrne, M\. Liakata, M\. Markham, P\. Pir, L\. N\. Soldatova, A\. Sparkes, K\. E\. Whelan, and A\. Clare\(2009\)The automation of science\.Science324\(5923\),pp\.85–89\.External Links:[Document](https://dx.doi.org/10.1126/science.1165620)Cited by:[§1](https://arxiv.org/html/2609.30971#S1.p1.1)\.
- \[12\]R\. Li, Z\. Hu, W\. Qu, J\. Zhang, Z\. Yin, S\. Zhang, X\. Huang, H\. Wang, T\. Wang, J\. Pang, W\. Ouyang, L\. Bai, W\. Zuo, L\. Duan, D\. Zhou, and S\. Tang\(2025\)LabUtopia: high\-fidelity simulation and hierarchical benchmark for scientific embodied agents\.InAdvances in Neural Information Processing Systems,Cited by:[§1](https://arxiv.org/html/2609.30971#S1.p2.1),[§1](https://arxiv.org/html/2609.30971#S1.p6.1.6.1.1),[§2\.1](https://arxiv.org/html/2609.30971#S2.SS1.p2.1),[§2\.2](https://arxiv.org/html/2609.30971#S2.SS2.p1.1)\.
- \[13\]S\. Li, Y\. Huang, C\. Guo, T\. Wu, J\. Zhang, L\. Zhang, and W\. Ding\(2025\)Chemistry3D: robotic interaction toolkit for chemistry experiments\.In2025 IEEE International Conference on Robotics and Automation \(ICRA\),pp\.8064–8071\.Cited by:[§1](https://arxiv.org/html/2609.30971#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.30971#S2.SS1.p2.1)\.
- \[14\]B\. Liu, Y\. Zhu, C\. Gao, Y\. Feng, Q\. Liu, Y\. Zhu, and P\. Stone\(2023\)LIBERO: benchmarking knowledge transfer for lifelong robot learning\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\.44776–44791\.Cited by:[§1](https://arxiv.org/html/2609.30971#S1.p6.1.4.1.1),[§2\.2](https://arxiv.org/html/2609.30971#S2.SS2.p1.1)\.
- \[15\]Z\. Liu, H\. Jin, Z\. Du, Z\. Wang, H\. Xu, P\. Li, J\. Gu, Q\. Lu, Q\. Wang, B\. Ji, and T\. Xiao\(2026\)An embodied simulation platform, benchmark, and data\-efficient augmentation framework for wet\-lab robotics\.External Links:2606\.12936Cited by:[§1](https://arxiv.org/html/2609.30971#S1.p2.1),[§1](https://arxiv.org/html/2609.30971#S1.p6.1.7.1.1),[§2\.1](https://arxiv.org/html/2609.30971#S2.SS1.p2.1)\.
- \[16\]O\. Mees, L\. Hermann, E\. Rosete\-Beas, and W\. Burgard\(2022\)CALVIN: a benchmark for language\-conditioned policy learning for long\-horizon robot manipulation tasks\.IEEE Robotics and Automation Letters7\(3\),pp\.7327–7334\.External Links:[Document](https://dx.doi.org/10.1109/LRA.2022.3180108)Cited by:[§2\.2](https://arxiv.org/html/2609.30971#S2.SS2.p1.1)\.
- \[17\]Physical Intelligence, K\. Black, N\. Brown, J\. Darpinian, K\. Dhabalia, D\. Driess, A\. Esmail, M\. Equi, C\. Finn, N\. Fusai, M\. Y\. Galliker, D\. Ghosh, L\. Groom, K\. Hausman, B\. Ichter, S\. Jakubczak, T\. Jones, L\. Ke, D\. LeBlanc, S\. Levine, A\. Li\-Bell, M\. Mothukuri, S\. Nair, K\. Pertsch, A\. Z\. Ren, L\. X\. Shi, L\. Smith, J\. T\. Springenberg, K\. Stachowicz, J\. Tanner, Q\. Vuong, H\. Walke, A\. Walling, H\. Wang, L\. Yu, and U\. Zhilinsky\(2025\)π0\.5\\pi\_\{0\.5\}: a vision\-language\-action model with open\-world generalization\.External Links:2504\.16054Cited by:[§5\.1\.1](https://arxiv.org/html/2609.30971#S5.SS1.SSS1.p1.1)\.
- \[18\]B\. Ren, X\. Liu, X\. Chen, Y\. Liu, C\. Li, D\. Gao, Z\. Su, J\. Xing, Z\. Xue, R\. Li, X\. Zhao, S\. Qiao, M\. Pan, W\. Zuo, L\. Bai, D\. Zhou, N\. Zhang, and H\. Chen\(2026\)LabVLA: grounding vision\-language\-action models in scientific laboratories\.External Links:2606\.13578Cited by:[§1](https://arxiv.org/html/2609.30971#S1.p2.1),[§1](https://arxiv.org/html/2609.30971#S1.p6.1.10.1.1),[§2\.2](https://arxiv.org/html/2609.30971#S2.SS2.p2.1)\.
- \[19\]M\. Savva, A\. Kadian, O\. Maksymets, Y\. Zhao, E\. Wijmans, B\. Jain, J\. Straub, J\. Liu, V\. Koltun, J\. Malik, D\. Parikh, and D\. Batra\(2019\)Habitat: a platform for embodied AI research\.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp\.9339–9347\.External Links:[Document](https://dx.doi.org/10.1109/ICCV.2019.00943)Cited by:[§1](https://arxiv.org/html/2609.30971#S1.p2.1)\.
- \[20\]B\. Shen, F\. Xia, C\. Li, R\. Martín\-Martín, L\. Fan, G\. Wang, C\. Pérez\-D’Arpino, S\. Buch, S\. Srivastava, L\. P\. Tchapmi, M\. E\. Tchapmi, K\. Vainio, J\. Wong, L\. Fei\-Fei, and S\. Savarese\(2021\)iGibson 1\.0: a simulation environment for interactive tasks in large realistic scenes\.In2021 IEEE/RSJ International Conference on Intelligent Robots and Systems,pp\.7520–7527\.External Links:[Document](https://dx.doi.org/10.1109/IROS51168.2021.9636667)Cited by:[§1](https://arxiv.org/html/2609.30971#S1.p2.1)\.
- \[21\]S\. Steiner, J\. Wolf, S\. Glatzel, A\. Andreou, J\. M\. Granda, G\. Keenan, T\. Hinkley, G\. Aragon\-Camarasa, P\. J\. Kitson, D\. Angelone, and L\. Cronin\(2019\)Organic synthesis in a modular robotic system driven by a chemical programming language\.Science363\(6423\),pp\.eaav2211\.External Links:[Document](https://dx.doi.org/10.1126/science.aav2211)Cited by:[§2\.1](https://arxiv.org/html/2609.30971#S2.SS1.p1.1)\.
- \[22\]N\. J\. Szymanski, B\. Rendy, Y\. Fei, R\. E\. Kumar, T\. He, D\. Milsted, M\. J\. McDermott, M\. Gallant, E\. D\. Cubuk, A\. Merchant, H\. Kim, A\. Jain, C\. J\. Bartel, K\. Persson, Y\. Zeng,et al\.\(2023\)An autonomous laboratory for the accelerated synthesis of inorganic materials\.Nature624\(7990\),pp\.86–91\.External Links:[Document](https://dx.doi.org/10.1038/s41586-023-06734-w)Cited by:[§1](https://arxiv.org/html/2609.30971#S1.p1.1)\.
- \[23\]E\. Todorov, T\. Erez, and Y\. Tassa\(2012\)MuJoCo: a physics engine for model\-based control\.In2012 IEEE/RSJ International Conference on Intelligent Robots and Systems,pp\.5026–5033\.External Links:[Document](https://dx.doi.org/10.1109/IROS.2012.6386109)Cited by:[§1](https://arxiv.org/html/2609.30971#S1.p2.1),[§3\.4](https://arxiv.org/html/2609.30971#S3.SS4.p1.1)\.
- \[24\]Y\. Wang, Z\. Xian, F\. Chen, T\. Wang, Y\. Wang, K\. Fragkiadaki, Z\. Erickson, D\. Held, and C\. Gan\(2024\)RoboGen: towards unleashing infinite data for automated robot learning via generative simulation\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\.51936–51983\.Cited by:[§1](https://arxiv.org/html/2609.30971#S1.p6.1.8.1.1),[§2\.2](https://arxiv.org/html/2609.30971#S2.SS2.p2.1)\.
- \[25\]S\. Zhang, Z\. Xu, P\. Liu, X\. Yu, Y\. Li, Q\. Gao, Z\. Fei, Z\. Yin, Z\. Wu, Y\. Jiang, and X\. Qiu\(2024\)VLABench: a large\-scale benchmark for language\-conditioned robotics manipulation with long\-horizon reasoning tasks\.External Links:2412\.18194Cited by:[§1](https://arxiv.org/html/2609.30971#S1.p6.1.5.1.1),[§2\.2](https://arxiv.org/html/2609.30971#S2.SS2.p1.1)\.
- \[26\]T\. Z\. Zhao, V\. Kumar, S\. Levine, and C\. Finn\(2023\)Learning fine\-grained bimanual manipulation with low\-cost hardware\.InProceedings of Robotics: Science and Systems,Daegu, Republic of Korea\.External Links:[Document](https://dx.doi.org/10.15607/RSS.2023.XIX.016)Cited by:[§5\.1\.1](https://arxiv.org/html/2609.30971#S5.SS1.SSS1.p1.1)\.
- \[27\]B\. Zitkovich, T\. Yu, S\. Xu, P\. Xu, T\. Xiao, F\. Xia, J\. Wu, P\. Wohlhart, S\. Welker, A\. Wahid,et al\.\(2023\)RT\-2: vision\-language\-action models transfer web knowledge to robotic control\.InProceedings of the 7th Conference on Robot Learning,Proceedings of Machine Learning Research, Vol\.229,pp\.2165–2183\.Cited by:[§1](https://arxiv.org/html/2609.30971#S1.p1.1),[§2\.2](https://arxiv.org/html/2609.30971#S2.SS2.p1.1)\.

相似文章

跨尺度科学挑战的AI智能体基准测试

arXiv cs.AI

介绍SciAgentArena,一个约200个任务的基准测试,用于评估真实科学研究中的AI智能体。发现智能体在明确指定的数据分析工作流程中表现有效,但在产生新颖见解和开放式探索方面存在困难。

从提示到协议:实验室自动化的AI代理

arXiv cs.AI

本文介绍了一种AI代理,它将大型语言模型与实验室编排软件集成,使科学家能够使用自然语言创建、监控和管理自动化的实验室协议。在三个模拟实验室上的评估显示,该代理实现了97%的首次尝试协议生成成功率,并且所需的界面操作大幅减少。