Embodied-BenchClaw: An Autonomous Multi-Agent System for Embodied Spatial Intelligence Benchmark Construction
Summary
This paper proposes Embodied-BenchClaw, an autonomous multi-agent system that automatically constructs embodied spatial intelligence benchmarks from user intent through a five-stage pipeline with process quality control and an extensible Skill Library.
View Cached Full Text
Cached at: 06/11/26, 01:49 PM
# Embodied-BenchClaw: An Autonomous Multi-Agent System for Embodied Spatial Intelligence Benchmark Construction
Source: [https://arxiv.org/html/2606.11909](https://arxiv.org/html/2606.11909)
Baoyang Jiang1Fengchun Zhang2Leyuan Wang3Haotian Li3Yida Wang3Zhe Ji4 Jinshan Lai2Xi Ren5Jianwei Hu1Qiang Ma1
1QiYuan Lab 2School of Information and Software Engineering, University of Electronic Science and Technology of China 3Beijing University of Posts and Telecommunications 4School of Computer Science and Engineering, Northeastern University 5School of Computer Science and Engineering, Beihang University
###### Abstract
Benchmarks are essential for evaluating and advancing embodied spatial intelligence, yet their construction remains labor\-intensive, hard to reuse, and difficult to maintain\. Existing embodied benchmarks are often released as static datasets and may quickly become saturated as model capabilities improve, reducing their ability to distinguish next\-generation models\. We proposeEmbodied\-BenchClaw, an autonomous agentic system for constructing embodied spatial intelligence benchmarks\. Given a user\-specified evaluation intent, Embodied\-BenchClaw automatically produces a complete and continually updatable benchmark package through a five\-stage pipeline, including intent blueprinting, data collection, structuring and cleaning, benchmark synthesis, and evaluation reporting\. The pipeline is coordinated by three agents for planning, construction, and evaluation, which respectively handle intent\-to\-blueprint conversion, benchmark generation, and evaluation\-driven refinement\. To improve reusability and reliability, Embodied\-BenchClaw introduces an extensible Skill Library that decomposes complex benchmark construction into composable, verifiable, and repairable building blocks, together with process quality control to constrain and correct low\-quality samples during generation\. We instantiate multiple Embodied\-BenchClaw\-produced benchmarks covering indoor spatial reasoning, outdoor spatial reasoning, robotic manipulation, quadruped robot navigation, UAV/aerial\-view understanding, and static benchmark enhancement\. These benchmarks span diverse embodied carriers, heterogeneous data sources, and fine\-grained spatial capabilities\. Extensive experiments, including human evaluation, judge\-based quality assessment, consistency checks, cost analysis, and ablation studies, show that Embodied\-BenchClaw can construct verifiable, executable, maintainable, and diagnostically useful embodied spatial benchmarks with reduced manual effort\. The results demonstrate its potential for scalable and continually refreshable evaluation of embodied spatial intelligence\.
Figure 1:Overview of Embodied\-BenchClaw\. Given a user request, Embodied\-BenchClaw recommends the benchmark scope, resources, capabilities, and evaluation plan, and then automatically constructs continually updatable benchmark packages through agentic construction and process quality control\. The construction process is supported by reusable resources, including skills, expert templates, capability/resource cards, simulators, static benchmarks, and real\-world images or videos\.## 1Introduction
Vision\-language models \(VLMs\), multimodal large language models \(MLLMs\), and vision\-language\-action models are progressing rapidly\. Beyond model scaling, recent research increasingly explores automated model improvement pipelines, including automatic data generation, self\-evaluation, agentic feedback, and iterative capability refinement\(Yuanet al\.,[2024](https://arxiv.org/html/2606.11909#bib.bib60); Gaoet al\.,[2025](https://arxiv.org/html/2606.11909#bib.bib61); Kielaet al\.,[2021](https://arxiv.org/html/2606.11909#bib.bib35)\)\. In this context, benchmarks are no longer only static leaderboards; they are key instruments for measuring emerging capabilities, exposing failure modes, and providing feedback for model evolution\(Kielaet al\.,[2021](https://arxiv.org/html/2606.11909#bib.bib35)\)\. However, benchmark construction remains much slower than model development\. High\-quality benchmarks are expensive to design, annotate, verify, and maintain, and once released, they may quickly become saturated, overfitted, or contaminated\(Kielaet al\.,[2021](https://arxiv.org/html/2606.11909#bib.bib35); Yanget al\.,[2023](https://arxiv.org/html/2606.11909#bib.bib62)\)\. This mismatch makes it difficult to timely evaluate new model capabilities and provide reliable feedback for model improvement\.
Figure 2:Representational similarity among embodied spatial intelligence benchmarks\. The similarity is computed from dataset\-level Qwen\-SAE activation fingerprintsDenget al\.\([2026](https://arxiv.org/html/2606.11909#bib.bib64)\), revealing potential redundancy and complementarity among existing benchmarks\.This bottleneck is especially prominent for embodied spatial intelligence\. Embodied agents require models to reason about egocentric spatial relations, object layouts, navigation cues, cross\-view correspondences, manipulation\-relevant constraints, and task feasibility from visual observations\. A large number of embodied AI benchmarks have been developed for navigation, instruction following, manipulation, 3D understanding, autonomous driving, robotic manipulation, and multimodal agent evaluation, as summarized in recent surveys\(Duanet al\.,[2022](https://arxiv.org/html/2606.11909#bib.bib57); Liuet al\.,[2025](https://arxiv.org/html/2606.11909#bib.bib58)\)\. Representative benchmarks such as ALFRED, LIBERO, EmbodiedScan, and EmbodiedBench provide valuable task definitions, data resources, and evaluation protocols\(Shridharet al\.,[2020](https://arxiv.org/html/2606.11909#bib.bib20); Liuet al\.,[2023](https://arxiv.org/html/2606.11909#bib.bib6); Wanget al\.,[2024](https://arxiv.org/html/2606.11909#bib.bib32); Yanget al\.,[2025](https://arxiv.org/html/2606.11909#bib.bib9)\)\. Many existing systems also automate parts of benchmark or data construction, such as procedural scene generation, template\-based task instantiation, demonstration synthesis, trajectory replay, simulator\-state verification, task generation, or reward generation\(Deitkeet al\.,[2022](https://arxiv.org/html/2606.11909#bib.bib38); Nasirianyet al\.,[2024](https://arxiv.org/html/2606.11909#bib.bib23); Mandlekaret al\.,[2023](https://arxiv.org/html/2606.11909#bib.bib40); Maet al\.,[2023](https://arxiv.org/html/2606.11909#bib.bib45)\)\. Nevertheless, most embodied benchmarks are still released as fixed artifacts tied to specific simulators, datasets, task definitions, or evaluation scripts\. When new VLMs, application scenarios, or spatial capabilities emerge, constructing a suitable benchmark still requires substantial manual effort in capability decomposition, resource selection, evidence checking, item design, scoring construction, and result analysis\.
Moreover, existing embodied spatial benchmarks may contain overlapping representational coverage\. As shown in Figure[2](https://arxiv.org/html/2606.11909#S1.F2), we analyze the representational similarity among existing embodied spatial intelligence benchmarks using Qwen\-SAE activation fingerprintsDenget al\.\([2026](https://arxiv.org/html/2606.11909#bib.bib64)\)\. For each benchmark, multimodal samples are mapped into sparse internal activation profiles, which are then averaged into dataset\-level fingerprints\. The pairwise cosine similarity heatmap shows that several benchmarks occupy highly similar regions in the model’s internal feature space, suggesting potential redundancy in current embodied spatial evaluation\. This observation motivates a construction framework that can not only generate new benchmarks, but also analyze, refine, and update benchmark packages toward under\-covered capability regions\.
To address this gap, we proposeEmbodied\-BenchClaw, a quality\-guided multi\-agent framework for automatically constructing VLM\-oriented embodied spatial benchmarks\. As shown in Figure[1](https://arxiv.org/html/2606.11909#S0.F1), Embodied\-BenchClaw transforms a high\-level user request into continually updatable benchmark packages\. It first interacts with the user to recommend and confirm the benchmark scope, resource selection, target capabilities, and evaluation plan\. It then performs automatic benchmark construction through three cooperating agents: a planning agent that converts user intents into construction plans, a construction agent that executes the five\-stage benchmark construction workflow, and a process quality\-control agent that verifies intermediate artifacts and triggers local repair when validation fails\. The construction process is supported by reusable resources, including a skill library, an expert template library, capability/resource cards, simulators, static benchmarks, and real\-world images or videos\.
Unlike one\-shot question generation, Embodied\-BenchClaw emphasizes resource\-aware construction, hierarchical skill\-guided execution, evidence\-grounded item synthesis, executable scoring, provenance tracing, and local repair\. Each construction stage is implemented through stage\-wise skill DAGs, where skills are executable units with input\-output contracts, tool bindings, validation rules, and optional repair actions\. The process quality\-control agent invokes executable verification to check intermediate artifacts, including schema validity, evidence completeness, answer derivability, scoring executability, response parsability, and trace completeness\. When verification fails, Embodied\-BenchClaw locates the affected construction step through provenance records and triggers local repair rather than restarting the full workflow\. The final output is not merely a set of isolated questions, but a task\-specific benchmark package containing items, metadata, evidence records, reference answers, scoring protocols, model response logs, evaluation reports, and update suggestions\. In the current implementation, Embodied\-BenchClaw focuses on image\-, multi\-view\-, and simulator\-state\-grounded embodied spatial evaluation, rather than closed\-loop robot control, tactile sensing, force feedback, or real\-time physical interaction\. By making benchmark construction reusable, verifiable, and maintainable, Embodied\-BenchClaw supports timely VLM evaluation, benchmark refinement, and lifecycle\-oriented benchmark maintenance\.
Our contributions are summarized as follows:
- •We proposeEmbodied\-BenchClaw, a fully automated multi\-agent framework for embodied spatial benchmark construction\. It consists of a planning agent, a construction agent, and a process quality\-control agent, enabling the transformation from user\-specified evaluation intents to continually updatable benchmark packages\.
- •We design a hierarchical skill\-guided execution mechanism with stage\-wise DAGs\. Each construction stage is executed through reusable skills with explicit input\-output contracts, executable components, validation rules, and repair actions, supporting modular and traceable benchmark construction\.
- •We introduce a process quality\-control mechanism based on executable verification, provenance tracing, and local repair\. This mechanism constrains agent behavior, checks intermediate artifacts, and improves the reliability and maintainability of benchmark construction\.
- •We instantiate six representative embodied spatial benchmarks, includingIndoor\-Bench,NavLand\-Bench,Drive\-Bench,Aerial\-Bench,Quad\-Bench, andManip\-Bench, covering indoor reasoning, indoor navigation, autonomous driving, UAV aerial perception, quadruped robot reasoning, and robotic manipulation\.
Table 1:Comparison with representative benchmark construction and evaluation systems\. “/yes” means the system explicitly supports the capability as part of its main design\. Multi\-type Data Sources refers to benchmark construction from multiple categories of data sources, such as simulators, real\-world visual data, existing benchmarks, and user\-provided data\.MethodModalityDomainEnd\-to\-endReusableEmbodiedMulti\-typeEvidence /FeedbackBenchmarkConstructionSkillsSpatialData SourcesScoring/ RepairOptimizationDynabench \(2021\)Single\-modalNLP✗✗✗✗✓✓✓WebArena \(2024\)Multi\-modalAgent✗✗✗✗✓✓✗RoboCasa \(2024\)Multi\-modalEmbodied AI✗✗✓✗✓✗✗RoboGen \(2024\)Multi\-modalEmbodied AI✗✗✓✗✓✓✗AutoBencher \(2025\)Single\-modalLLM Evaluation✓✗✗✗✓✓✓BenchAgents \(2025\)Single\-modalLLM Evaluation✓✗✗✗✓✓✗EmbodiedBench \(2025\)Multi\-modalEmbodied AI✗✗✓✗✓✗✗Code2Bench \(2026\)Multi\-modalSoftware Engineering✓✗✗✗✓✓✓Embodied\-BenchClaw \(Ours\)Multi\-modalEmbodied AI✓✓✓✓✓✓✓
## 2Related Work
Automated benchmark construction has recently emerged as an important direction for evaluating rapidly evolving models\. Dynabench studies human\-and\-model\-in\-the\-loop dynamic benchmarking\(Kielaet al\.,[2021](https://arxiv.org/html/2606.11909#bib.bib35)\), while AutoBencher, BenchBench, and BenchAgents automate benchmark creation for LLM or multimodal capabilities\(Liet al\.,[2024](https://arxiv.org/html/2606.11909#bib.bib3); Zhenget al\.,[2026](https://arxiv.org/html/2606.11909#bib.bib4); Buttet al\.,[2024](https://arxiv.org/html/2606.11909#bib.bib63)\)\. BenchAgents is particularly related to our work: it decomposes benchmark creation into planning, generation, data verification, and evaluation agents, and uses human\-in\-the\-loop feedback to control benchmark diversity and quality\(Buttet al\.,[2024](https://arxiv.org/html/2606.11909#bib.bib63)\)\. However, its main focus is general generative capabilities such as planning and constraint satisfaction during text generation, rather than embodied spatial benchmark construction\. Other systems focus on executable or environment\-grounded evaluation\. Code2Bench and PRDBench construct code\-oriented benchmarks with executable tests\(Zhanget al\.,[2026b](https://arxiv.org/html/2606.11909#bib.bib36); Fuet al\.,[2025](https://arxiv.org/html/2606.11909#bib.bib37)\); WebArena, SWE\-bench, and Claw\-Eval\-Live evaluate agents in web, software, or workflow environments\(Zhouet al\.,[2024](https://arxiv.org/html/2606.11909#bib.bib52); Jimenezet al\.,[2024](https://arxiv.org/html/2606.11909#bib.bib53); Liet al\.,[2026](https://arxiv.org/html/2606.11909#bib.bib56)\)\. A2Eval is closer to benchmark optimization, as it reorganizes existing embodied VLM benchmarks for more compact and efficient evaluation\(Zhanget al\.,[2026a](https://arxiv.org/html/2606.11909#bib.bib1)\)\. Table[1](https://arxiv.org/html/2606.11909#S1.T1)summarizes the comparison\.
## 3Embodied\-BenchClaw
### 3\.1Overview
Embodied\-BenchClaw is an automated benchmark construction framework for VLM\-oriented embodied spatial intelligence\. Given a high\-level user request, Embodied\-BenchClaw first interacts with the user to clarify the target capability scope, recommended data sources, evaluation settings, and model roster\. After the user confirms the construction plan, Embodied\-BenchClaw invokes reusable skills, expert templates, capability cards, and quality\-feedback mechanisms to construct a complete benchmark package\.
The output of Embodied\-BenchClaw is not a set of isolated questions\. Instead, it is a benchmark package containing benchmark items, metadata, evidence records, reference answers, scoring protocols, model responses, leaderboards, diagnostic reports, and update suggestions\. This design allows the constructed benchmark to support both immediate VLM evaluation and later benchmark refinement\.
Figure[3](https://arxiv.org/html/2606.11909#S3.F3)illustrates the overall workflow\. Embodied\-BenchClaw organizes benchmark construction into five automated stages, while the reusable skill library provides executable construction operations across these stages\. Quality gates are enforced at stage boundaries to detect construction errors, trace their causes, and trigger local repair when necessary\. The following subsections first summarize the five\-stage workflow, then describe the skill library and quality\-feedback mechanism in detail, followed by the expert templates and capability cards that provide reusable construction knowledge\.
Figure 3:Skill\-driven benchmark construction workflow with quality feedback\. Embodied\-BenchClaw invokes stage\-specific skills to transform user benchmark requests into task\-specific benchmark packages through five stages: intent blueprinting, data collection, structuring and cleaning, benchmark synthesis, and evaluation reporting\. Quality gates, failure diagnosis, provenance tracing, and local repair enable targeted feedback to earlier stages\.
### 3\.2Automated Construction Workflow
Embodied\-BenchClaw organizes benchmark construction into five stages\. Each stage invokes a group of reusable skills and produces intermediate artifacts for the next stage\. This staged design makes the construction process modular, auditable, and repairable, while keeping the overall workflow fully automated after user confirmation\.
#### Stage 1: Intent Blueprinting\.
This stage converts a rough benchmark request into a structured benchmark blueprint, including target capabilities, task scope, candidate resources, and preliminary evaluation requirements\. It invokes planning\-related skills to clarify the benchmark intent and prepare an executable construction plan\.
#### Stage 2: Data Collection\.
This stage collects candidate data according to the blueprint, using simulators, real\-world visual data, user\-provided data, or existing benchmarks when available\. It invokes acquisition\-related skills to register data sources and collect raw samples with traceable source records\.
#### Stage 3: Structuring and Cleaning\.
This stage converts heterogeneous raw data into cleaned and structured evidence records\. It invokes evidence\-related skills to normalize formats, prepare annotations, remove invalid samples, and organize metadata for later benchmark synthesis\.
#### Stage 4: Benchmark Synthesis\.
This stage materializes benchmark items from expert templates and verified evidence\. It invokes synthesis\-related skills to bind templates with evidence, generate questions and reference answers, create scoring logic, and package valid benchmark items\.
#### Stage 5: Evaluation Reporting\.
This stage evaluates target VLMs on the constructed benchmark package and generates diagnostic reports\. It invokes evaluation\-related skills to run model APIs, parse responses, compute scores, build leaderboards, and produce capability\-wise and source\-wise analyses\.
### 3\.3Hierarchical Skill\-Guided Execution with Stage\-wise DAGs
Embodied\-BenchClaw follows an agentic execution paradigm, where a construction agent advances benchmark construction by invoking reusable skills within a fixed five\-stage pipeline\. The agent does not freely generate the workflow structure\. Instead, each construction stage is implemented as a predefined stage\-wise skill DAG, whose nodes are skills and whose edges represent artifact dependencies\. This design makes the construction process controllable, reproducible, and traceable, while still allowing the agent to execute different skills according to the current artifacts, data sources, and construction requirements\.
#### Skill definition\.
In Embodied\-BenchClaw, a skill is not a prompt alone\. It is a contract\-defined execution unit that specifies its input artifacts, output artifacts, executable component, validation rule, and optional repair action\. The executable component can be an LLM call, a Python script, a simulator adapter, an annotation tool, a scoring function, a data cleaning operator, or a validator\. Thus, Python code is treated as part of a skill when it implements a construction, verification, or repair operation under a declared input\-output contract\. This definition separates skill execution from free\-form prompting: prompts are mainly used for intent understanding, language realization, and report summarization, while evidence extraction, answer derivation, scoring, and formal checks are handled by tools, programs, and validators\.
Formally, each skill is represented as
𝒮k=⟨ℐk,𝒪k,ℰk,𝒱k,ℛk,τk⟩,\\mathcal\{S\}\_\{k\}=\\langle\\mathcal\{I\}\_\{k\},\\mathcal\{O\}\_\{k\},\\mathcal\{E\}\_\{k\},\\mathcal\{V\}\_\{k\},\\mathcal\{R\}\_\{k\},\\tau\_\{k\}\\rangle,whereℐk\\mathcal\{I\}\_\{k\}and𝒪k\\mathcal\{O\}\_\{k\}denote the input and output artifact schemas,ℰk\\mathcal\{E\}\_\{k\}denotes the executable component,𝒱k\\mathcal\{V\}\_\{k\}denotes the validation rule,ℛk\\mathcal\{R\}\_\{k\}denotes the optional repair action, andτk\\tau\_\{k\}denotes stage, capability, evidence, or tool tags used for skill selection\. A skill can be executed only when its required input artifacts are available and its preconditions are satisfied\.
#### Hierarchical skill roles\.
To support the full benchmark construction process, Embodied\-BenchClaw groups skills into four roles\.*Control skills*manage stage transitions, artifact dependency tracking, DAG execution, and provenance logging\.*Construction skills*produce benchmark artifacts, including collected data, structured evidence, benchmark items, reference answers, scoring scripts, and evaluation reports\.*Verification skills*implement process quality control through executable checks, such as schema validation, evidence completeness checking, answer derivability testing, ambiguity checking, duplication detection, and scoring executability testing\.*Repair skills*are triggered when verification fails; they localize the responsible upstream node, revise invalid artifacts, and rerun only the affected subgraph\. This role\-based hierarchy makes quality control and feedback repair part of skill\-guided execution rather than an external post\-processing step\.
#### Stage\-wise skill DAGs\.
For each construction stagess, Embodied\-BenchClaw uses a predefined skill DAG
𝒢s=\(𝒩s,𝒜s\),\\mathcal\{G\}\_\{s\}=\(\\mathcal\{N\}\_\{s\},\\mathcal\{A\}\_\{s\}\),where𝒩s\\mathcal\{N\}\_\{s\}denotes the skill nodes in stagessand𝒜s\\mathcal\{A\}\_\{s\}denotes artifact\-dependency edges\. A node consumes input artifacts, executes its corresponding skill, and produces output artifacts for downstream nodes\. Independent nodes can be executed in parallel, while dependent nodes are scheduled according to the DAG order\. Because the DAG is predefined by developers rather than generated freely by the LLM, Embodied\-BenchClaw avoids unstable workflow planning while still allowing different skills to be activated according to the benchmark request and available resources\.
Verification and repair are also represented as skill nodes in the stage\-wise DAG\. After key construction nodes produce artifacts, verification nodes check whether these artifacts satisfy schema constraints, evidence requirements, answer derivation rules, scoring constraints, and traceability requirements\. When a check fails, the construction agent uses the artifact dependency graph and provenance records to locate the responsible node, invokes the corresponding repair skill, and reruns the affected subgraph\. This local repair strategy avoids restarting the full five\-stage workflow when only a local artifact is invalid\.
Table[2](https://arxiv.org/html/2606.11909#S3.T2)summarizes how hierarchical skill\-guided execution is instantiated across the five construction stages\. The table highlights the predefined DAG nodes in each stage, the corresponding skill roles, the artifact contracts, and the verification or repair focus\.
Table 2:Hierarchical skill\-guided execution in Embodied\-BenchClaw\. Each construction stage is implemented as a predefined skill DAG consisting of control, construction, verification, and repair skills\. Quality control is embedded through verification and repair nodes rather than applied only as final post\-processing\.Stage / LoopRepresentative DAG NodesSkill RolesArtifact ContractVerification / Repair Focus\\rowcolor\[HTML\]F6F9FE Stage 1: Intent BlueprintingIntent parsing; capability decomposition; resource recommendation; metric drafting; workflow planning; blueprint validation\.Control; Construction; VerificationUser request→\\rightarrowintent blueprint, capability scope, resource plan, metric draft, execution plan\.Checks intent consistency, capability–resource feasibility, metric alignment, and unresolved ambiguities\.\\rowcolor\[HTML\]F5FAF5 Stage 2: Data CollectionSource registration; simulator acquisition; real\-data acquisition; user\-data ingestion; existing\-benchmark ingestion; source validation\.Construction; Verification; RepairResource plan→\\rightarrowraw samples, source records, preliminary metadata, collection log\.Checks source availability, schema validity, collection completeness, and source traceability; repairs missing or invalid source records\.\\rowcolor\[HTML\]F6FAFC Stage 3: Structuring and CleaningEvidence normalization; annotation preparation; simulator\-GT conversion; tool invocation; data cleaning; evidence validation\.Construction; Verification; RepairRaw samples→\\rightarrowunified evidence pool, annotations, cleaned records, metadata, tool logs\.Checks evidence completeness, GT traceability, annotation validity, visibility constraints, and cleaning validity; repairs invalid or incomplete evidence records\.\\rowcolor\[HTML\]FDF9F4 Stage 4: Benchmark SynthesisTemplate–evidence binding; question expansion; answer\-program execution; item filtering; scoring\-script generation; artifact packaging\.Construction; Verification; RepairEvidence pool \+ templates→\\rightarrowbenchmark items, reference answers, evidence links, scoring scripts, benchmark package\.Checks evidence–template compatibility, answer derivability, scoring executability, ambiguity, duplication, and package completeness; repairs invalid bindings or regenerated items\.\\rowcolor\[HTML\]FAF9FD Stage 5: Evaluation ReportingModel evaluation; response parsing; score aggregation; leaderboard construction; capability\-wise analysis; report validation\.Construction; Verification; RepairBenchmark package \+ model outputs→\\rightarrowparsed responses, scores, leaderboard, evaluation report\.Checks raw\-output preservation, parsing validity, score reproducibility, abnormal responses, and report completeness; repairs parsing or reporting errors\.\\rowcolor\[HTML\]FDF8FB Cross\-stage Feedback and RepairGate aggregation; provenance tracing; defect localization; affected\-subgraph rerun; release validation\.Control; Verification; RepairFailed artifact \+ execution trace→\\rightarrowdefect report, repair plan, updated artifacts, release decision\.Localizes the responsible skill node or stage, triggers targeted repair, reruns affected subgraphs, and produces PASS / REVIEW / FAIL decisions\.
This skill\-guided execution design provides the main interface for extending Embodied\-BenchClaw\. When a new simulator, dataset, annotation tool, model API, or scoring protocol is introduced, developers can add or update the corresponding skill implementation under the same input\-output contract\. Embodied\-BenchClaw records skill\-level execution traces, including invoked skills, consumed and generated artifacts, validation results, failure causes, repair actions, runtime, and token cost\. These traces support debugging, auditing, efficiency analysis, and later refinement of construction skills\.
### 3\.4Contract\-guided Process Verification and Repair
Process quality control in Embodied\-BenchClaw is implemented as contract\-guided verification over stage\-wise skill DAGs, rather than as a final post\-processing step\. Inspired by verification\-driven agent workflows, Embodied\-BenchClaw first converts the user request and confirmed construction plan into a structured construction contract\. The contract specifies the target capabilities, allowed resources, required evidence fields, expected artifacts, scoring constraints, and stage transition conditions\. It serves as the shared source of truth for the construction agent, preventing semantic drift across stages and skill calls\.
Each skill node is associated with an input\-output contract\. Formally, a skill is defined as
𝒮k=⟨ℐk,𝒪k,𝒫k,𝒬k,ℰk,ℛk⟩,\\mathcal\{S\}\_\{k\}=\\langle\\mathcal\{I\}\_\{k\},\\mathcal\{O\}\_\{k\},\\mathcal\{P\}\_\{k\},\\mathcal\{Q\}\_\{k\},\\mathcal\{E\}\_\{k\},\\mathcal\{R\}\_\{k\}\\rangle,whereℐk\\mathcal\{I\}\_\{k\}and𝒪k\\mathcal\{O\}\_\{k\}are the input and output artifact schemas,𝒫k\\mathcal\{P\}\_\{k\}and𝒬k\\mathcal\{Q\}\_\{k\}are the precondition and postcondition,ℰk\\mathcal\{E\}\_\{k\}is the executable component, andℛk\\mathcal\{R\}\_\{k\}is the repair action\. For each stagess, Embodied\-BenchClaw executes a predefined skill DAG𝒢s=\(𝒩s,𝒜s\)\\mathcal\{G\}\_\{s\}=\(\\mathcal\{N\}\_\{s\},\\mathcal\{A\}\_\{s\}\), where nodes are skill executions and edges are artifact dependencies\.
For each skill nodevkv\_\{k\}, Embodied\-BenchClaw applies a process verifier
Φ\(vk\)=ϕsafe\(vk\)∧ϕdag\(vk\)∧ϕcontract\(ok\)∧ϕquality\(ok\),\\Phi\(v\_\{k\}\)=\\phi\_\{\\mathrm\{safe\}\}\(v\_\{k\}\)\\land\\phi\_\{\\mathrm\{dag\}\}\(v\_\{k\}\)\\land\\phi\_\{\\mathrm\{contract\}\}\(o\_\{k\}\)\\land\\phi\_\{\\mathrm\{quality\}\}\(o\_\{k\}\),whereϕsafe\\phi\_\{\\mathrm\{safe\}\}checks whether the agent uses only allowed skills, tools, files, and resources;ϕdag\\phi\_\{\\mathrm\{dag\}\}checks whether the execution follows the predefined stage DAG and satisfies all input dependencies;ϕcontract\\phi\_\{\\mathrm\{contract\}\}checks schema validity and artifact consistency; andϕquality\\phi\_\{\\mathrm\{quality\}\}checks benchmark\-specific requirements such as evidence completeness, answer derivability, scoring executability, ambiguity, duplication, response parsability, and report completeness\. Only artifacts that pass these verification conditions are allowed to flow to downstream nodes\.
At the stage boundary, Embodied\-BenchClaw aggregates node\-level verification results through a quality gate:
Qs=∏vk∈𝒩s𝟏\[Φ\(vk\)=1\]\.Q\_\{s\}=\\prod\_\{v\_\{k\}\\in\\mathcal\{N\}\_\{s\}\}\\mathbf\{1\}\\left\[\\Phi\(v\_\{k\}\)=1\\right\]\.IfQs=1Q\_\{s\}=1, the workflow proceeds to the next stage\. Otherwise, Embodied\-BenchClaw identifies the failed node set
ℱs=\{vk∈𝒩s∣Φ\(vk\)=0\}\\mathcal\{F\}\_\{s\}=\\\{v\_\{k\}\\in\\mathcal\{N\}\_\{s\}\\mid\\Phi\(v\_\{k\}\)=0\\\}and uses provenance traces to locate the affected subgraph
ℋs=Desc\(ℱs;𝒢s\)\.\\mathcal\{H\}\_\{s\}=\\mathrm\{Desc\}\(\\mathcal\{F\}\_\{s\};\\mathcal\{G\}\_\{s\}\)\.Repair is then performed only onℋs\\mathcal\{H\}\_\{s\}rather than on the entire workflow\. Typical repair actions include recollecting missing data, rerunning annotation tools, revising template–evidence bindings, regenerating invalid items, repairing scoring scripts, or rerunning response parsing\. The repaired artifacts are re\-verified before being released to downstream stages\.
This mechanism addresses two requirements of agentic benchmark construction\. First, it improves agent safety by restricting skill execution to declared resources, tools, and stage\-DAG dependencies\. Second, it constrains agent behavior toward benchmark quality by ensuring that every generated item is supported by traceable evidence, has a derivable reference answer, and can be scored by an executable protocol\. Therefore, process quality control in Embodied\-BenchClaw is not merely manual inspection or LLM self\-critique, but a combination of contract checking, executable verification, provenance tracing, and local repair\.
### 3\.5Expert Templates and Capability Cards
Besides executable skills, Embodied\-BenchClaw maintains two types of reusable knowledge resources: expert templates and capability cards\. Expert templates define task patterns for embodied spatial evaluation\. Each template specifies the target capability, question form, required evidence fields, valid conditions, answer space, and scoring constraints\. During benchmark synthesis, templates are not directly used as final questions; they are instantiated only when the required evidence is available and passes the corresponding quality gate\.
Capability cards describe the resources that Embodied\-BenchClaw can use during construction, including simulators, datasets, user\-provided data formats, annotation tools, cleaning operators, and model APIs\. A capability card records what a resource can provide, how it can be invoked, what input and output formats it supports, and what constraints or failure modes should be considered\. During intent blueprinting and data collection, Embodied\-BenchClaw uses capability cards to match user requests with available resources and avoid constructing benchmark items that cannot be supported by evidence\.
Together, expert templates and capability cards provide the structured knowledge required by the skill library\. Skills execute construction operations, while templates and cards specify what should be constructed and which resources can support it\. This separation allows Embodied\-BenchClaw to extend its construction capability by adding new skills, templates, or capability cards without rewriting the full benchmark construction workflow\.
## 4Embodied\-BenchClaw\-produced Benchmarks
Embodied\-BenchClaw \(E\-BenchClaw\) produces standardized benchmark packages for embodied spatial intelligence evaluation\. Each package is constructed from a user\-specified evaluation intent and contains benchmark items, input observations, reference answers, evidence records, scoring protocols, metadata, provenance traces, and update suggestions\. Unlike conventional static benchmarks that are usually tied to a single data source or task format, E\-BenchClaw supports both*new benchmark construction*and*existing benchmark enhancement*\. The former builds new evaluation sets from heterogeneous embodied resources, while the latter reuses images or scenes from saturated benchmarks and generates more fine\-grained, evidence\-dependent spatial reasoning questions\.
To cover representative embodied settings, we construct six types of E\-BenchClaw\-produced benchmarks: indoor spatial reasoning, outdoor spatial reasoning, object manipulation, quadruped robot navigation, UAV/aerial\-view understanding, and static benchmark enhancement\. These benchmarks cover diverse data sources, embodied agents, observation forms, spatial skills, and action constraints\. This design allows us to evaluate whether E\-BenchClaw can generalize beyond a single scenario and support different embodied carriers, including egocentric agents, autonomous vehicles, robotic arms, quadruped robots, and UAVs\.
Table 3:Overview of E\-BenchClaw\-produced benchmarks\. Each benchmark package contains items, observations, evidence records, reference answers, scoring protocols, metadata, provenance traces, and update suggestions\.Benchmark TypePrimary EmbodimentTypical Data SourcesInput FormsCore Spatial SkillsKey ConstraintsIndoor Spatial ReasoningEgocentric agents and indoor mobile robotsIndoor images/videos, RGB\-D data, simulated indoor scenes, and indoor QA/navigation resourcesImages, videos, depth maps, semantic maps, room layouts, and target descriptionsObject localization, directional relations, occlusion reasoning, room topology, target reachability, and path selectionIndoor obstacles, room connectivity, field\-of\-view limits, and traversable areasOutdoor Spatial ReasoningAutonomous vehicles and ground mobile platformsStreet\-view data, road videos, inspection data, open\-environment images, trajectories, and mapsOutdoor images/videos, road regions, obstacles, target areas, trajectories, and mapsDrivable\-area recognition, obstacle reasoning, target localization, distance\-scale reasoning, occlusion reasoning, and path\-risk assessmentRoad boundaries, physical obstacles, safety constraints, and dynamic risksObject ManipulationRobotic armsTabletop scenes, robotic\-arm simulation, real\-world manipulation images/videos, and detection/segmentation resultsRGB/RGB\-D images, object masks, object poses, grasp points, and task instructionsGraspability, support and containment relations, placement feasibility, operation ordering, and affordance reasoningReachable workspace, collision constraints, placement stability, and container capacityQuadruped Robot NavigationQuadruped robots and ground inspection robotsRough terrains, stairs, slopes, narrow passages, simulated environments, and real robot\-view dataEgocentric images/videos, terrain labels, obstacles, trajectories, and target positionsTerrain traversability, detour/crossing decisions, passage selection, local path planning, and stability\-risk judgmentObstacle height, slope, stair size, passage width, and locomotion limitsUAV / Aerial\-view UnderstandingUAVs and aerial inspection platformsAerial images/videos, aerial inspection data, target\-search scenes, and 3D terrain or map dataAerial images/videos, target boxes, region annotations, flight altitude, flight routes, and mapsAerial localization, region membership, global\-local matching, height/distance reasoning, visibility range, and flight\-path feasibilityFlight altitude, field of view, occlusion, no\-fly zones, and route constraintsStatic Benchmark EnhancementThe original embodiment or visual input format of the source benchmarkExisting embodied or VLM benchmarks where current models already achieve high scoresOriginal images/scenes are kept unchanged; questions, answers, evidence, and scoring rules are reconstructedFine\-grained spatial reasoning, multi\-step relation integration, counterfactual viewpoint reasoning, occlusion reasoning, and path/action feasibility judgmentThe original data remain unchanged, while questions are made more dependent on spatial evidence and deeper reasoning
For each benchmark type, E\-BenchClaw outputs a unified benchmark package rather than isolated QA samples\. The evidence records may include image regions, object masks, depth cues, semantic maps, trajectories, camera poses, traversable regions, action constraints, or simulator states\. The reference answers are accompanied by derivation traces, indicating how they can be obtained from visual\-spatial evidence, geometric rules, simulator metadata, or constraint checking\. This standardized output format makes the generated benchmarks inspectable, traceable, and executable for automatic evaluation\.
## 5Experiments
We evaluate E\-BenchClaw from two perspectives\. First, we evaluate it as an automated benchmark construction framework, focusing on whether it can reliably convert user requests into complete, evidence\-grounded, and executable benchmark packages\. Second, we evaluate the generated benchmarks, focusing on whether they provide meaningful signals for comparing VLMs and MLLMs on embodied spatial intelligence\.
### 5\.1Experimental Setup
#### Research questions\.
Our experiments are organized around four questions:RQ1asks whether E\-BenchClaw can construct benchmarks across diverse embodied scenarios and data sources\.RQ2asks whether the generated items are reliable, evidence\-grounded, spatially consistent, and executable for scoring\.RQ3asks whether E\-BenchClaw improves construction efficiency and whether its core components are necessary\.RQ4asks whether the generated benchmarks provide useful diagnostic signals for evaluating embodied spatial capabilities of VLMs/MLLMs\.
#### Construction configuration\.
Unless otherwise specified, E\-BenchClaw uses an qwen3\.6\-35b\-a3b LLM backbone to coordinate the planning, construction, and quality\-control agents\. All construction runs follow the same pipeline: intent blueprinting, data collection, structuring and cleaning, benchmark synthesis, and evaluation reporting\. The system uses a reusable Skill Library for data processing, evidence extraction, spatial verification, scoring, and repair; an Expert Template Library for embodied spatial task generation; and process quality control for executable verification and local repair\.
#### Evaluation dimensions\.
We summarize the evaluation dimensions in Table[4](https://arxiv.org/html/2606.11909#S5.T4)\. The experiments jointly evaluate coverage, reliability, controllability, diagnostic value, and efficiency, while avoiding reliance on a single quality signal\.
Table 4:Main evaluation dimensions for E\-BenchClaw\.DimensionEvaluation FocusMain MetricsPurposeScenario and Embodiment CoverageWhether E\-BenchClaw supports diverse scenes and embodied agentsNumber of benchmark types, number of data sources, embodiment coverage, skill coverageTo verify that the framework is not limited to a single embodied taskConstruction ReliabilityWhether the generated packages are complete, valid, and executableConstruction success rate, package completeness, valid yield rate, scoring\-script success rateTo verify reliable end\-to\-end benchmark constructionEvidence\-grounded QualityWhether items are supported by traceable visual\-spatial evidenceEvidence binding rate, evidence\-grounding correctness, answer derivability, human correctnessTo verify that answers are not unsupported LLM guessesSpatial and Action ConsistencyWhether spatial relations and actions satisfy geometric and physical constraintsSpatial consistency pass rate, action feasibility pass rate, GT executability rateTo verify that generated items respect embodied constraintsProcess ControllabilityWhether low\-quality artifacts can be detected and repaired during constructionQuality\-gate pass rates, ambiguity rate, duplicate rate, repair success rateTo verify the effectiveness of process quality control and local repairEvaluation ValueWhether the generated benchmarks can distinguish model capabilitiesOverall accuracy, skill\-wise scores, difficulty consistency, model\-scale consistencyTo verify model separability and diagnostic usefulnessEfficiencyWhether E\-BenchClaw reduces human effort and supports fast updatesTime per valid sample, human review time, token cost, tool calls, update timeTo verify scalability and maintainability
#### Human evaluation and judge\-based evaluation\.
We conduct human evaluation on sampled items\. Annotators inspect the input observation, question, reference answer, evidence record, and answer derivation trace, and rate each item in terms of correctness, clarity, answerability, and relevance to the target embodied spatial skill\. Scores are assigned on a 1–5 scale and normalized to 0–100\. We report the human acceptance rate as the percentage of samples whose average score exceeds a predefined threshold\.
We also use LLM/VLM\-as\-Judge as a scalable complementary evaluation\. At the benchmark level, we report User\-Intention Alignment \(UIA\)\. At the item level, we report Format and Schema Quality \(FSQ\), Question–Answer Coherence \(QAC\), Context–Question Correspondence \(CQC\), Target Skill Dependency \(TSD\), and Skill\-Specific Challenge \(SSC\)\. For embodied benchmarks, we additionally report Evidence\-Grounding Correctness \(EGC\), Spatial\-Geometric Consistency \(SGC\), Action Feasibility Consistency \(AFC\), and Ground\-Truth Executability \(GTE\)\.
#### Automatic quality gates\.
During construction, E\-BenchClaw applies automatic quality gates to intermediate and final artifacts\. These gates check format validity, evidence binding, GT executability, spatial consistency, action feasibility, ambiguity, redundancy, and scoring executability\. Invalid samples are either locally repaired or discarded\. We report the pass rate of each quality gate, the overall valid yield rate, and the repair success rate\.
#### Comparison methods\.
We compare Full E\-BenchClaw with four construction baselines:Direct LLM Generation, which directly prompts an LLM to generate benchmark items;Template\-only Generation, which instantiates questions from fixed templates;LLM \+ Template Generation, which uses an LLM to select and fill templates but removes executable verification; andHuman\-assisted Construction, which relies on human\-written or human\-revised items as a high\-quality reference setting\. This comparison evaluates whether E\-BenchClaw improves both construction quality and efficiency over simpler alternatives\.
#### Ablation studies\.
We ablate the major components of E\-BenchClaw to measure their contributions\.w/o Automated Planningremoves explicit intent decomposition and construction planning\.w/o Skill Libraryremoves reusable construction and verification skills\.w/o Expert Template Libraryreplaces embodied expert templates with generic QA templates\.w/o Formal Verificationdisables geometric, physical, and GT\-executability checks\.w/o Process Quality Controlremoves stage\-wise quality gates\.w/o Local Repairdetects invalid artifacts but does not repair them\.w/o Deduplicationremoves redundancy and coverage control\. We evaluate these variants using human acceptance, judge\-based scores, valid yield rate, repair success rate, construction cost, and model evaluation results\.
#### Model evaluation\.
To test benchmark utility, we evaluate a set of closed\-source and open\-source VLMs/MLLMs on the generated benchmarks\. For multiple\-choice and short\-answer items, we report accuracy or exact match\. For grounding tasks, we report grounding accuracy or IoU when annotations are available\. For navigation and manipulation\-related items, we additionally report route validity, collision\-free rate, reachability accuracy, affordance accuracy, and action\-feasibility accuracy when executable checks are available\. We further report skill\-wise scores to analyze model weaknesses across direction, distance, occlusion, reachability, path planning, object relations, affordance, and multi\-view reasoning\.
#### Efficiency analysis\.
We report time per valid sample, human review time, token cost, number of tool calls, valid yield rate, repair rate, repair success rate, and benchmark update time\. These metrics measure whether E\-BenchClaw can reduce manual effort, improve valid sample production, and support continual benchmark refresh\.
### 5\.2Results on UAV / Aerial\-view Understanding Benchmark
We analyze one representative E\-BenchClaw\-produced benchmark for UAV/aerial\-view spatial understanding\. This benchmark contains 5,000 generated questions and covers five capability dimensions: visible object/category recognition, counting with interval choices, image\-plane spatial relations, visible\-area and salient\-scale comparison, and depth\-based near\-far reasoning\. The benchmark is instantiated into five question types: single\-choice selection, multi\-choice selection, binary comparison, interval selection, and ranking\.
#### Overall results\.
Table[5](https://arxiv.org/html/2606.11909#S5.T5)reports the average performance under three settings: vision\-language evaluation, blind evaluation, and random guessing\. The average score of vision\-language models reaches65\.39%65\.39\\%, while the blind setting drops to29\.82%29\.82\\%, which is close to the random baseline of33\.30%33\.30\\%\. This large gap indicates that the benchmark cannot be solved mainly by language priors or answer\-option bias; instead, models need to rely on aerial visual\-spatial evidence\.
Table 5:Average performance on the UAV/aerial\-view benchmark under different evaluation settings\.SettingOverallSingleMultiBinaryIntervalRankingVision\-language Avg\.65\.3974\.7952\.2984\.3632\.0783\.46Blind Avg\.29\.8230\.0018\.0043\.4317\.2140\.45Random33\.3028\.008\.5056\.0028\.0046\.00VL \- Blind Gap\+35\.57\+44\.79\+34\.29\+40\.93\+14\.86\+43\.01
#### Complete vision\-language results\.
Table[6](https://arxiv.org/html/2606.11909#S5.T6)reports the complete results of all evaluated models under the vision\-language setting\. GPT\-5\.5 achieves the best overall score of74\.13%74\.13\\%, followed by Gemini\-3\-Pro\-Preview, Kimi\-K2\.5, Qwen3\.6\-27B, and Qwen3\.6\-35B\-A3B\. The results show that the benchmark provides meaningful model separability: strong models generally perform better, but different models exhibit different strengths across question types\. For example, Kimi\-K2\.5 obtains the highest binary\-comparison score, Qwen3\.6\-27B performs best on interval selection, and GPT\-5\.5 achieves the strongest overall and ranking performance\.
Table 6:Complete model performance on the UAV/aerial\-view benchmark under the vision\-language setting\. Scores are reported in percentage\. Models are grouped into closed\-source/API models and open\-source/local models\.ModelOverallSingleMultiBinaryIntervalRankingClosed\-source / API ModelsGPT\-5\.574\.1389\.0054\.0093\.0042\.0092\.67Gemini\-3\-Pro\-Preview71\.8390\.0054\.0093\.0033\.0089\.17Kimi\-K2\.571\.8083\.0056\.5095\.0035\.0089\.50Grok\-4\.20\-0309\-Reasoning68\.5779\.0054\.5091\.0033\.0085\.33Moonshot\-VL\-128K\-Vision\-Preview67\.5382\.0053\.0084\.0032\.0086\.67Claude\-Opus\-4\-767\.2382\.0056\.0083\.0030\.0085\.17Claude\-Sonnet\-4\-667\.0779\.0055\.0085\.0031\.0085\.33GPT\-5\.265\.8774\.0053\.0085\.0032\.0085\.33GPT\-4\.163\.1368\.0063\.0079\.0024\.0081\.67Qwen3\-VL\-Plus56\.8761\.0044\.0078\.0023\.0078\.33Claude\-Haiku\-4\-5\-2025100151\.3355\.0043\.5067\.0024\.0067\.17Qwen3\-VL\-Flash\-2025\-10\-1549\.8750\.0037\.5067\.0023\.0071\.83Open\-source / Local ModelsQwen3\.6\-27B71\.2777\.0055\.0094\.0044\.0086\.33Qwen3\.6\-35B\-A3B69\.0078\.0053\.0087\.0043\.0084\.00
#### Complete blind and random results\.
Table[5](https://arxiv.org/html/2606.11909#S5.T5)reports the complete blind\-evaluation results and the random baseline\. In the blind setting, the average score is only29\.82%29\.82\\%, much lower than the vision\-language average of65\.39%65\.39\\%\. This confirms that the benchmark strongly depends on visual evidence from aerial observations\. Notably, several blind results are close to or below the random baseline, indicating that language\-only guessing is insufficient for this benchmark\.
#### Question\-type difficulty\.
The benchmark reveals a clear difficulty hierarchy\. Binary comparison and ranking are relatively easier, with vision\-language averages of84\.36%84\.36\\%and83\.46%83\.46\\%, respectively\. Single\-choice questions are also tractable, reaching74\.79%74\.79\\%\. In contrast, multi\-choice questions are much harder, with an average score of52\.29%52\.29\\%, because they require set\-level reasoning over multiple visible targets\. Interval selection is the most challenging type, with an average score of only32\.07%32\.07\\%\. This suggests that aerial counting, quantity\-range estimation, and depth/scale\-related interval reasoning remain difficult for current VLMs/MLLMs\.
#### Implications\.
Overall, the UAV/aerial\-view benchmark demonstrates three important properties of E\-BenchClaw\-produced benchmarks\. First, the large gap between vision\-language and blind evaluation confirms strong evidence dependency\. Second, the performance differences across models show that the benchmark has meaningful discriminative power\. Third, the large variation across question types provides fine\-grained diagnostic signals, especially revealing weaknesses in multi\-object selection and interval\-based aerial spatial reasoning\.
## References
- Benchagents: automated benchmark creation with agent interaction\.InICLR 2025 Workshop on Navigating and Addressing Data Problems for Foundation Models,Cited by:[§2](https://arxiv.org/html/2606.11909#S2.p1.1)\.
- M\. Deitke, E\. VanderBilt, A\. Herrasti, L\. Weihs, K\. Ehsani, J\. Salvador, W\. Han, E\. Kolve, A\. Kembhavi, and R\. Mottaghi \(2022\)ProcTHOR: large\-scale embodied ai using procedural generation\.Advances in Neural Information Processing Systems35,pp\. 5982–5994\.Cited by:[§1](https://arxiv.org/html/2606.11909#S1.p2.1)\.
- B\. Deng, X\. Wang, Y\. Wang, Y\. Wan, Y\. Ma, B\. Yang, H\. Wei, J\. Tang, H\. Lin, R\. Gao, T\. Li, Q\. Cao, X\. Ren, X\. Deng, A\. Yang, F\. Huang, D\. Liu, and J\. Zhou \(2026\)Qwen\-Scope: turning sparse features into development tools for large language models\.External Links:2605\.11887,[Link](https://arxiv.org/abs/2605.11887)Cited by:[Figure 2](https://arxiv.org/html/2606.11909#S1.F2),[§1](https://arxiv.org/html/2606.11909#S1.p3.1)\.
- J\. Duan, S\. Yu, H\. L\. Tan, H\. Zhu, and C\. Tan \(2022\)A survey of embodied ai: from simulators to research tasks\.IEEE Transactions on Emerging Topics in Computational Intelligence6\(2\),pp\. 230–244\.Cited by:[§1](https://arxiv.org/html/2606.11909#S1.p2.1)\.
- L\. Fu, B\. Zhang, H\. Guan, Y\. Zhu, L\. Qiu, W\. Liu, X\. Cao, X\. Cai, W\. Zhang, and Y\. Yu \(2025\)Automatically benchmarking llm code agents through agent\-driven annotation and evaluation\.arXiv preprint arXiv:2510\.24358\.Cited by:[§2](https://arxiv.org/html/2606.11909#S2.p1.1)\.
- H\. Gao, J\. Geng, W\. Hua, M\. Hu, X\. Juan, H\. Liu, S\. Liu, J\. Qiu, X\. Qi, Y\. Wu,et al\.\(2025\)A survey of self\-evolving agents: on path to artificial super intelligence\.arXiv preprint arXiv:2507\.210461\.Cited by:[§1](https://arxiv.org/html/2606.11909#S1.p1.1)\.
- C\. E\. Jimenez, J\. Yang, A\. Wettig, S\. Yao, K\. Pei, O\. Press, and K\. Narasimhan \(2024\)Swe\-bench: can language models resolve real\-world github issues?\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 54107–54157\.Cited by:[§2](https://arxiv.org/html/2606.11909#S2.p1.1)\.
- D\. Kiela, M\. Bartolo, Y\. Nie, D\. Kaushik, A\. Geiger, Z\. Wu, B\. Vidgen, G\. Prasad, A\. Singh, P\. Ringshia,et al\.\(2021\)Dynabench: rethinking benchmarking in nlp\.InProceedings of the 2021 conference of the North American chapter of the Association for Computational Linguistics: human language technologies,pp\. 4110–4124\.Cited by:[§1](https://arxiv.org/html/2606.11909#S1.p1.1),[§2](https://arxiv.org/html/2606.11909#S2.p1.1)\.
- C\. Li, Z\. Tang, H\. Lin, Y\. Lin, S\. Huang, S\. Liu, B\. Ye, R\. Li, L\. Li, B\. Wang,et al\.\(2026\)Claw\-eval\-live: a live agent benchmark for evolving real\-world workflows\.arXiv preprint arXiv:2604\.28139\.Cited by:[§2](https://arxiv.org/html/2606.11909#S2.p1.1)\.
- X\. L\. Li, F\. Kaiyom, E\. Z\. Liu, Y\. Mai, P\. Liang, and T\. Hashimoto \(2024\)Autobencher: towards declarative benchmark construction\.arXiv preprint arXiv:2407\.08351\.Cited by:[§2](https://arxiv.org/html/2606.11909#S2.p1.1)\.
- B\. Liu, Y\. Zhu, C\. Gao, Y\. Feng, Q\. Liu, Y\. Zhu, and P\. Stone \(2023\)Libero: benchmarking knowledge transfer for lifelong robot learning\.Advances in Neural Information Processing Systems36,pp\. 44776–44791\.Cited by:[§1](https://arxiv.org/html/2606.11909#S1.p2.1)\.
- Y\. Liu, W\. Chen, Y\. Bai, X\. Liang, G\. Li, W\. Gao, and L\. Lin \(2025\)Aligning cyber space with physical world: a comprehensive survey on embodied ai\.IEEE/ASME Transactions on Mechatronics\.Cited by:[§1](https://arxiv.org/html/2606.11909#S1.p2.1)\.
- Y\. J\. Ma, W\. Liang, G\. Wang, D\. Huang, O\. Bastani, D\. Jayaraman, Y\. Zhu, L\. Fan, and A\. Anandkumar \(2023\)Eureka: human\-level reward design via coding large language models\.arXiv preprint arXiv:2310\.12931\.Cited by:[§1](https://arxiv.org/html/2606.11909#S1.p2.1)\.
- A\. Mandlekar, S\. Nasiriany, B\. Wen, I\. Akinola, Y\. Narang, L\. Fan, Y\. Zhu, and D\. Fox \(2023\)Mimicgen: a data generation system for scalable robot learning using human demonstrations\.arXiv preprint arXiv:2310\.17596\.Cited by:[§1](https://arxiv.org/html/2606.11909#S1.p2.1)\.
- S\. Nasiriany, A\. Maddukuri, L\. Zhang, A\. Parikh, A\. Lo, A\. Joshi, A\. Mandlekar, and Y\. Zhu \(2024\)Robocasa: large\-scale simulation of everyday tasks for generalist robots\.arXiv preprint arXiv:2406\.02523\.Cited by:[§1](https://arxiv.org/html/2606.11909#S1.p2.1)\.
- M\. Shridhar, J\. Thomason, D\. Gordon, Y\. Bisk, W\. Han, R\. Mottaghi, L\. Zettlemoyer, and D\. Fox \(2020\)Alfred: a benchmark for interpreting grounded instructions for everyday tasks\.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition,pp\. 10740–10749\.Cited by:[§1](https://arxiv.org/html/2606.11909#S1.p2.1)\.
- T\. Wang, X\. Mao, C\. Zhu, R\. Xu, R\. Lyu, P\. Li, X\. Chen, W\. Zhang, K\. Chen, T\. Xue,et al\.\(2024\)Embodiedscan: a holistic multi\-modal 3d perception suite towards embodied ai\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 19757–19767\.Cited by:[§1](https://arxiv.org/html/2606.11909#S1.p2.1)\.
- R\. Yang, H\. Chen, J\. Zhang, M\. Zhao, C\. Qian, K\. Wang, Q\. Wang, T\. V\. Koripella, M\. Movahedi, M\. Li,et al\.\(2025\)Embodiedbench: comprehensive benchmarking multi\-modal large language models for vision\-driven embodied agents\.arXiv preprint arXiv:2502\.09560\.Cited by:[§1](https://arxiv.org/html/2606.11909#S1.p2.1)\.
- S\. Yang, W\. Chiang, L\. Zheng, J\. E\. Gonzalez, and I\. Stoica \(2023\)Rethinking benchmark and contamination for language models with rephrased samples\.arXiv preprint arXiv:2311\.04850\.Cited by:[§1](https://arxiv.org/html/2606.11909#S1.p1.1)\.
- W\. Yuan, R\. Y\. Pang, K\. Cho, X\. Li, S\. Sukhbaatar, J\. Xu, and J\. Weston \(2024\)Self\-rewarding language models\.arXiv preprint arXiv:2401\.10020\.Cited by:[§1](https://arxiv.org/html/2606.11909#S1.p1.1)\.
- S\. Zhang, J\. Hu, Z\. Chen, Z\. Ding, Y\. Zhang, Y\. Zhang, Z\. Zhou, J\. Liao, S\. Zhou, Y\. Dai,et al\.\(2026a\)A2Eval: agentic and automated evaluation for embodied brain\.arXiv preprint arXiv:2602\.01640\.Cited by:[§2](https://arxiv.org/html/2606.11909#S2.p1.1)\.
- Z\. Zhang, R\. Liu, A\. Liu, X\. Liu, X\. Gao, and H\. Sun \(2026b\)Code2Bench: scaling source and rigor for dynamic benchmark construction\.InThe Fourteenth International Conference on Learning Representations,Cited by:[§2](https://arxiv.org/html/2606.11909#S2.p1.1)\.
- Y\. Zheng, H\. Luo, Z\. Lin, W\. Liu, and L\. A\. Tuan \(2026\)BenchBench: benchmarking automated benchmark generation\.arXiv preprint arXiv:2603\.20807\.Cited by:[§2](https://arxiv.org/html/2606.11909#S2.p1.1)\.
- S\. Zhou, F\. F\. Xu, H\. Zhu, X\. Zhou, R\. Lo, A\. Sridhar, X\. Cheng, T\. Ou, Y\. Bisk, D\. Fried,et al\.\(2024\)Webarena: a realistic web environment for building autonomous agents\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 15585–15606\.Cited by:[§2](https://arxiv.org/html/2606.11909#S2.p1.1)\.Similar Articles
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop
Introduces ESI-BENCH, a comprehensive benchmark for embodied spatial intelligence built on OmniGibson, covering 10 task categories and 29 subcategories. Experiments show active exploration substantially outperforms passive approaches, with failures mainly due to action blindness rather than perception, revealing a metacognitive gap in models compared to humans.
Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence
This paper introduces SIS-Bench, a benchmark for evaluating self-awareness and spatial cognition in UAV embodied intelligence using multimodal large language models, and explores motion-aware representations to improve performance.
VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control
VABench introduces a benchmark to evaluate embodied spatial intelligence in models by testing their ability to observe, reason, and act through visual demonstrations and active perception. It shows that active camera control improves task success, but no model completes long-horizon episodes.
Embodied Operators and Benchmarking: Toward Reusable and Deployable Embodied Intelligence Systems
This paper defines embodied operators as reusable functional modules for embodied intelligence pipelines, presents a taxonomy covering five major categories, and proposes a multi-dimensional benchmark framework for evaluating their deployability and composability.
SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning
SpatialClaw is a training-free framework that uses code as an action interface to enable flexible, stateful spatial reasoning in vision-language models, achieving superior performance across diverse 3D/4D spatial reasoning tasks.