MetaSpace: Metamorphic Testing for Spatial Cognition in Embodied Agents

arXiv cs.AI Papers

Summary

MetaSpace is a framework that applies metamorphic testing to evaluate spatial cognition in embodied agents, generating test cases from execution trajectories and encoding logical/physical constraints as Prolog rules. Benchmarking shows state-of-the-art MLLM-driven agents score far below human levels on spatial cognition, with directional tasks being especially weak.

arXiv:2608.07533v1 Announce Type: new Abstract: An embodied agent is an intelligent entity that interacts with its environment through a physical body. Currently, the evaluation of embodied agents primarily relies on two paradigms: (1) manually annotated Visual Question Answering (VQA) pairs and (2) high-level task completion metrics, such as success in navigation or manipulation. The former is labor-intensive and subject to variability in annotation quality. The latter may obscure critical vulnerabilities, allowing agents to complete tasks through suboptimal means or safety violations, thereby concealing safety risks and inefficiencies. Given that spatial cognition is the cornerstone for executing embodied tasks, there is a pressing need to assess whether embodied agents possess robust spatial cognition during task execution. Inspired by metamorphic testing principles in software engineering, we propose MetaSpace, a novel framework designed to evaluate the spatial cognition of agents. By leveraging spatiotemporal multimodal states derived from real execution trajectories, MetaSpace automatically generates test cases based on predefined metamorphic relations (MRs) grounded in logical rules and physical laws. Crucially, we encode these MRs as executable rules in a logic programming language (Prolog). Violations of these relations indicate failures in spatial cognition. Our empirical evaluation across three embodied scenarios demonstrates that MetaSpace successfully detects 90,422 spatial cognition errors in state-of-the-art (SOTA) MLLM-driven agents. We introduce the Spatial Cognition (SC) score to quantify performance. Results indicate that all SOTA agents achieve average scores between 0.44 and 0.52, significantly lower than the human benchmark of 0.96.
Original Article
View Cached Full Text

Cached at: 08/11/26, 08:02 AM

# MetaSpace: Metamorphic Testing for Spatial Cognition in Embodied Agents
Source: [https://arxiv.org/html/2608.07533](https://arxiv.org/html/2608.07533)
\(2026\-02\-17\)

###### Abstract\.

An embodied agent is an intelligent entity that interacts with its environment through a physical body\. Currently, the evaluation of embodied agents primarily relies on two paradigms: \(1\) manually annotated Visual Question Answering \(VQA\) pairs and \(2\) high\-level task completion metrics, such as success in navigation or manipulation\. The former is labor\-intensive and subject to variability in annotation quality\. The latter may obscure critical vulnerabilities, allowing agents to complete tasks through suboptimal means or safety violations, thereby concealing safety risks and inefficiencies\. Given that spatial cognition is the cornerstone for executing embodied tasks, there is a pressing need to assess whether embodied agents possess robust spatial cognition during task execution\.

Inspired by metamorphic testing principles in software engineering, we propose MetaSpace, a novel framework designed to evaluate the spatial cognition of agents\. By leveraging spatiotemporal multimodal states derived from real execution trajectories, MetaSpace automatically generates test cases based on predefined metamorphic relations \(MRs\) grounded in logical rules and physical laws\. Crucially, we encode these MRs as executable rules in a logic programming language \(Prolog\)\. Violations of these relations indicate failures in spatial cognition\. Our empirical evaluation across three embodied scenarios demonstrates that MetaSpace successfully detects 90,422 spatial cognition errors in state\-of\-the\-art \(SOTA\) MLLM\-driven agents\. We introduce the Spatial Cognition \(SC\) score to quantify performance\. Results indicate that all SOTA agents achieve average scores between 0\.44 and 0\.52, significantly lower than the human benchmark of 0\.96\. Additionally, these agents struggle with directional tasks, with SC scores consistently below 0\.38\. In contrast, their performance in magnitude\-related tasks is relatively better, with most SC scores exceeding 0\.5\. To mitigate the identified spatial cognition errors, we explore potential improvement strategies\. Preliminary results suggest that traditional prompting techniques \(e\.g\., Chain of Thought\) are limited, while spatially\-aware prompting \(e\.g\., cognitive maps\) shows promise\. Our findings underscore the importance of ongoing community efforts to enhance embodied agent performance by prioritizing the improvement of spatial cognition, a fundamental requirement for executing embodied tasks\.

embodied agent, spatial cognition, software testing

††copyright:cc††doi:10\.1145/3798212††journalyear:2026††journal:PACMPL††journalvolume:10††journalnumber:OOPSLA1††article:104††publicationmonth:4††ccs:Software and its engineering Software testing and debugging## 1\.Introduction

Embodied agents are attracting significant attention from both academia and industry, with applications spanning robotics\(Roy et al\.,[2021](https://arxiv.org/html/2608.07533#bib.bib52); Yang et al\.,[2025](https://arxiv.org/html/2608.07533#bib.bib70); Li et al\.,[2024b](https://arxiv.org/html/2608.07533#bib.bib28)\), autonomous driving\(Tian et al\.,[2025](https://arxiv.org/html/2608.07533#bib.bib55)\), and drones\(Zhao et al\.,[2025](https://arxiv.org/html/2608.07533#bib.bib79); Gao et al\.,[2024](https://arxiv.org/html/2608.07533#bib.bib20)\)\. Embodied agents are capable of performing a wide range of tasks, from semantic tasks \(e\.g\., household chores\) to core embodied functions \(e\.g\., navigation, rearrangement, and manipulation\)\. Leveraging Multi\-modal Large Language Models \(MLLMs\) to create embodied agents presents a promising avenue for tackling embodied tasks\(Li et al\.,[2024b](https://arxiv.org/html/2608.07533#bib.bib28); Yang et al\.,[2025](https://arxiv.org/html/2608.07533#bib.bib70)\)\. However, while Large Language Models \(LLMs\) have achieved remarkable success in linguistic tasks\(Rostam et al\.,[2024](https://arxiv.org/html/2608.07533#bib.bib51); Naveed et al\.,[2025](https://arxiv.org/html/2608.07533#bib.bib40); Raiaan et al\.,[2024](https://arxiv.org/html/2608.07533#bib.bib47); Dong et al\.,[2025](https://arxiv.org/html/2608.07533#bib.bib17)\), they face critical challenges in visuospatial tasks, which significantly undermine their effectiveness and reliability in embodied intelligence applications\.

To evaluate the performance of embodied agents, the standard paradigm in the AI and robotics community relies on benchmarking with manually annotated Visual Question Answering \(VQA\) pairs \(e\.g\., Multiple\-Choice Questions \(MCQs\)\)\(Du et al\.,[2024](https://arxiv.org/html/2608.07533#bib.bib18); Ramakrishnan et al\.,[2025](https://arxiv.org/html/2608.07533#bib.bib48); Ma et al\.,[2025](https://arxiv.org/html/2608.07533#bib.bib34); Zhao et al\.,[2025](https://arxiv.org/html/2608.07533#bib.bib79); Cheng et al\.,[2025](https://arxiv.org/html/2608.07533#bib.bib13); Gao et al\.,[2024](https://arxiv.org/html/2608.07533#bib.bib20); Dang et al\.,[2025](https://arxiv.org/html/2608.07533#bib.bib16)\)\. However, the manual design and annotation of test cases are labor\-intensive, and variability in annotator expertise can introduce inconsistency and bias into benchmark assessments\. Additionally, some benchmarks \(e\.g\.,\(Ramakrishnan et al\.,[2025](https://arxiv.org/html/2608.07533#bib.bib48)\)\) utilize classic cognitive psychology questions originally designed for humans or animals, such as the Minnesota Paper Form Board \(MPFB\) test\(Likert and Quasha,[1941](https://arxiv.org/html/2608.07533#bib.bib31)\)\. These approaches fail to capture the essence of*embodiment*in agents, leading to inadequate reflection of their actual performance in real\-world applications\.

![Refer to caption](https://arxiv.org/html/2608.07533v1/x1.png)\(a\)Wrong Magnitude Perception
![Refer to caption](https://arxiv.org/html/2608.07533v1/x2.png)\(b\)Redundant Trial\-and\-Error
![Refer to caption](https://arxiv.org/html/2608.07533v1/x3.png)\(c\)Coincidence\-Driven Success

Figure 1\.Examples of “false positive success” in high\-level embodied tasks\.In contrast, some studies have begun to rely on high\-level task completion metrics, such as whether an agent reaches navigation targets or successfully performs manipulation tasks\(Yang et al\.,[2025](https://arxiv.org/html/2608.07533#bib.bib70); Choi et al\.,[2024](https://arxiv.org/html/2608.07533#bib.bib14); Zhou et al\.,[2024](https://arxiv.org/html/2608.07533#bib.bib81); Cheng et al\.,[2025](https://arxiv.org/html/2608.07533#bib.bib13)\)\.111In this context, we define low\-level tasks as the foundational tasks that underpin high\-level tasks \(e\.g\., navigation and manipulation\)\. For instance, successful high\-level navigation relies on low\-level reasoning about object directional relations to avoid obstacles\.While intuitive and easy to quantify, these outcome\-oriented evaluations conceal underlying flaws, and agents may succeed in completing tasks through non\-optimal means or safety violations \(“false positive success”\)\. For example, \(1\)Wrong magnitude perception: as shown in Fig\.[1\(a\)](https://arxiv.org/html/2608.07533#S1.F1.sf1), agents may complete tasks based on erroneous action magnitude perception \(e\.g\., taking three steps for a navigation move, when only two are necessary\)\. In virtual environments, the absence of physical collisions allows agents to create an illusion of success\. \(2\)Redundant trial\-and\-error: agents may misjudge spatial information \(e\.g\., direction perception errors\) and correct their trajectories through repeated adjustments, thereby sacrificing efficiency for task completion without resulting in task failure \(Fig\.[1\(b\)](https://arxiv.org/html/2608.07533#S1.F1.sf2)\)\. \(3\)Coincidence\-driven success: spatial cognition errors may not always result in failure, as agents might, by chance, avoid negative consequences \(e\.g\., in Fig\.[1\(c\)](https://arxiv.org/html/2608.07533#S1.F1.sf3), the agent misjudges a 2\-meter door as1\.8meters1\.8\\text\{\\,\}\\mathrm\{m\}\\mathrm\{e\}\\mathrm\{t\}\\mathrm\{e\}\\mathrm\{r\}\\mathrm\{s\}but successfully passes through without getting stuck due to its own height of1\.7meters1\.7\\text\{\\,\}\\mathrm\{m\}\\mathrm\{e\}\\mathrm\{t\}\\mathrm\{e\}\\mathrm\{r\}\\mathrm\{s\}\)\. These “false positive successes” are not rare; existing embodied benchmark studies have similar observations through manual checking\(Li et al\.,[2024b](https://arxiv.org/html/2608.07533#bib.bib28); Yang et al\.,[2025](https://arxiv.org/html/2608.07533#bib.bib70); Cheng et al\.,[2025](https://arxiv.org/html/2608.07533#bib.bib13)\)\. Such defects in the existing evaluation paradigm conceal significant safety risks \(e\.g\., non\-catastrophic collisions\) and operational inefficiencies \(e\.g\., suboptimal paths\), undermining the trustworthiness of embodied agents in real\-world applications\.

The high\-level task outcome\-oriented evaluation paradigm is an end\-to\-end approach that assesses only the final results of task execution, without examining the intermediate processes or decision\-making involved\. In reality, embodied task execution is inherently compositional: successful completion of high\-level tasks relies on first accomplishing a series of spatial cognitive sub\-tasks\. For example, to navigate to a target location, an agent must understand the spatial relationships between the target and surrounding landmarks, and accurately perceive its own movement direction and distance to ensure correct actions\. In other words, spatial cognition \(e\.g\., movement perception, spatial awareness\) serves as the cornerstone for embodied agents to complete high\-level embodied tasks\. These observations highlight our key motivation for this research:

> “Is spatial cognition, as the keystone of high\-level embodied tasks, truly a solved problem for embodied agents?”

To bridge the identified research gap, it is essential to move beyond outcome\-oriented testing paradigms and adopt a capability\-oriented approach that enables an authentic evaluation of embodied capabilities\. Developing such a capability\-oriented evaluation framework requires us to address three significant challenges:C1: Inadequacy of Static Evaluation for Embodiment\.Existing methods for spatial cognition evaluations often fail to capture the dynamic essence of embodiment\. Most rely on static Visual Question Answering \(VQA\) tasks, which lack interactive engagement with the environment and do not assess the agent’s ability to transform between egocentric and allocentric perspectives\.C2: Challenge in Defining Test Oracle\.Automatically determining the ground truth of spatial cognition is inherently challenging\. Embodied tasks are dynamic and context\-dependent, making it difficult to define the expected spatial relations and perceptions for all possible scenarios\.C3: Challenge in Test Case Generation\.Manual design and annotation of test cases continues to be the prevailing practice\. However, this approach is both labor\-intensive and prone to incompleteness, frequently missing rare or edge cases for robust testing\. Furthermore, the quality and consistency of benchmark questions can vary significantly based on the expertise of human annotators, introducing noise and potential bias into the assessment\.

To address the above challenges, we propose MetaSpace, a metamorphic testing \(MT\)\(Chen et al\.,[2020](https://arxiv.org/html/2608.07533#bib.bib12)\)framework specifically designed to evaluate the spatial cognition of embodied agents\. MetaSpace shifts the focus of embodied agent assessment from external outcome to internal capability, specifically targeting SC abilities\. Importantly, all test cases are automatically generated from real embodied task execution trajectories, ensuring a strong correlation between evaluation results and actual embodied performance\. To addressC1andC3, MetaSpace collects spatiotemporal multimodal states derived from real trajectories and utilizes these states to generate test cases\. To addressC2, we craft a set of metamorphic relations \(MRs\) to serve as test oracles; these MRs are based on principled logical rules \(e\.g\., transitivity, symmetry\) and physical laws \(e\.g\., perspective geometry\)\. We use logic programming to encode these MRs as executable rules and agents’ spatial observations as facts to automate the validation process\. MetaSpace provides a comprehensive assessment of eight key embodied spatial cognitive abilities, ensuring a robust evaluation of embodied agents in real\-world embodied scenarios\. In summary, our contributions are threefold:

- •At the conceptual level, we move beyond simple end\-to\-end outcomes and instead scrutinize the intrinsic spatial cognitive abilities of embodied agents\. This capability\-based view unlocks transparency into the decision process itself, which is the foundation for high\-level embodied tasks\. We identify eight key spatial cognitive abilities essential for embodied agents, drawing insights from cognitive psychology and embodied intelligence literature\.
- •At the technical level, we develop MetaSpace, a MT framework that implements a set of logic\- and physics\-based MRs to evaluate embodied spatial cognition\. These MRs are encoded into a Prolog knowledge base, allowing for scalable and oracle\-free validation\.
- •At the empirical level, we apply MetaSpace to test six state\-of\-the\-art \(SOTA\) MLLM\-driven embodied agents across three real\-world scenarios\. MetaSpace executes 30,300 unique test cases on six agents, uncovering a total of 90,422 instances that trigger spatial cognitive errors\. Based on our findings, we derive valuable insights and offer recommendations for mitigating spatial cognition errors in embodied agents\. Preliminary experiments demonstrate that cognitive map prompting can enhance agents’ directional spatial cognitive abilities\.

## 2\.Background

### 2\.1\.Embodied Spatial Cognition

Definition\.The concept of spatial cognition originates from cognitive psychology\. It refers to the study of knowledge and beliefs regarding the spatial properties of objects and events, including aspects such as location, size, distance, and movement\(Montello,[2001](https://arxiv.org/html/2608.07533#bib.bib39)\)\. We particularly focus onembodied spatial cognition, which pertains to the spatial cognition of embodied intelligence in supporting embodied tasks\.222In this paper, we use “spatial cognition”, “SC”, and “spatial cognitive capability” interchangeably to describe “embodied spatial cognition”\.

Scope\.While humans can derive spatial cognition from various modalities \(e\.g\., one can estimate the location of a ringing phone even without seeing it\), most embodied agents primarily rely on visual input at this stage\(Burgess,[2008](https://arxiv.org/html/2608.07533#bib.bib7)\)\. Given this context, our work focuses on visual\-spatial cognition\. Additionally, while cognitive psychology provides classic spatial cognition experiments, including pen\-and\-paper tasks and the Minnesota Paper Form Board Test \(MPFB\)\(Likert and Quasha,[1941](https://arxiv.org/html/2608.07533#bib.bib31)\), our emphasis is on spatial cognition within embodied scenarios, which involves dynamic interactions with the environment\.

![Refer to caption](https://arxiv.org/html/2608.07533v1/x4.png)Figure 2\.Taxonomy of embodied spatial cognition with numbering codes\.Taxonomy\.We present a taxonomy of capabilities essential for embodied spatial cognition \(Fig\.[2](https://arxiv.org/html/2608.07533#S2.F2)\)\. Rather than an arbitrary collection, our taxonomy synthesizes foundational domains from cognitive science and psychology\. Specifically, we identify four key aspects, each grounded in established theories:SC1: Movement perception\.Grounded in Gibson’s ecological theory of perception\(Gibson,[2014](https://arxiv.org/html/2608.07533#bib.bib21)\), this capability captures the agent’s ability to sense self\-motion for immediate control\. It is foundational for basic navigation and manipulation tasks\.SC2: Spatial reasoning\.Drawing on Kosslyn’s theory of spatial relations\(Kosslyn,[1987](https://arxiv.org/html/2608.07533#bib.bib27)\), we assess the ability to identify spatial relationships between objects and landmarks\. Furthermore, based on Burgess’s model of spatial memory\(Burgess,[2006](https://arxiv.org/html/2608.07533#bib.bib6)\),SC4: egocentric\-allocentric transformationassesses the critical ability to translate between egocentric views and allocentric mental maps\. This is vital for envisioning actions from future viewpoints\.SC3: Perspective visualization\.This is essential as embodied agents operate in 3D environments\. Aligned with Marr’s computational vision theory\(Marr,[2010](https://arxiv.org/html/2608.07533#bib.bib37)\), we assess the recovery of intrinsic properties \(i\.e\., size and depth\) from 2D observations \(e\.g\., the 2\.5D sketch\)\. Across SC1, SC2, and SC4, the distinction between directional \(SC\*\-a\) and magnitude\-based \(SC\*\-b\) capabilities is supported by Kosslyn’s theory of spatial relations\(Kosslyn,[1987](https://arxiv.org/html/2608.07533#bib.bib27)\), which differentiates between categorical spatial processing \(e\.g\., relative directions\) and coordinate spatial processing \(e\.g\., precise distances\)\. This suggests they involve distinct cognitive mechanisms, warranting separate evaluation\. In contrast, SC3 focuses on recovering intrinsic properties \(size and depth\) from 2D observations, consistent with Marr’s vision theory\(Marr,[2010](https://arxiv.org/html/2608.07533#bib.bib37)\)\.

### 2\.2\.Logic Programming for Spatial Cognition Validation

In this study, we apply logic programming to implement the designed MRs for evaluating embodied spatial cognition\. In the context of embodied spatial cognition, we encode spatial observations as facts and MRs as rules, facilitating rigorous validation of the cognitive consistency of embodied agents without human intervention\. We elaborate on the specific process below\.

Automatic Fact Generation\.Agent responses regarding spatial observations are automatically converted into Prolog facts\.333MetaSpace constrains agents to respond in predefined formats through structured prompting\. We discuss the influence of structured prompting in[Fig\.13](https://arxiv.org/html/2608.07533#S6.F13)\.For example, when an agent perceives movement from states1s\_\{1\}tos2s\_\{2\}as “going forward”, this generates the fact𝑚𝑜𝑣𝑒𝐹𝑜𝑟𝑤𝑎𝑟𝑑​\(s1,s2\)\.\\mathit\{moveForward\(s\_\{1\},s\_\{2\}\)\.\}Similarly, the object relationships observed by agents can be encoded into facts such as𝑒𝑎𝑠𝑡𝑂𝑓​\(𝑙𝑎𝑛𝑑𝑚𝑎𝑟𝑘A,𝑙𝑎𝑛𝑑𝑚𝑎𝑟𝑘B\)\.\\mathit\{eastOf\(landmark\_\{A\},landmark\_\{B\}\)\.\}

Metamorphic Relation Rules\.A rule is a conditional statement that allows new facts to be inferred from existing ones\. A common structure for a rule is the Horn clause, which comprises a head predicate and a rule body \(a list of predicates\)\. In MetaSpace, each MR is encoded as Horn clause rules defining consistency constraints\. An example demonstrating the transitivity rule is𝑒𝑎𝑠𝑡𝑂𝑓​\(X,Z\)​:–​𝑒𝑎𝑠𝑡𝑂𝑓​\(X,Y\),𝑒𝑎𝑠𝑡𝑂𝑓​\(Y,Z\)\\mathit\{eastOf\(X,Z\)\\,\\text\{:\-\-\}\\,eastOf\(X,Y\),eastOf\(Y,Z\)\}\. This rule means that if X is east of Y, and Y is east of Z, the system can infer that X is east of Z\. Another example,𝑤𝑒𝑠𝑡𝑂𝑓​\(X,Y\)​:–​𝑒𝑎𝑠𝑡𝑂𝑓​\(Y,X\)\\mathit\{westOf\(X,Y\)\\,\\text\{:\-\-\}\\,eastOf\(Y,X\)\}, defines𝑤𝑒𝑠𝑡𝑂𝑓\\mathit\{westOf\}as an inverse relation of𝑒𝑎𝑠𝑡𝑂𝑓\\mathit\{eastOf\}\.

Automated MR Validation Program\.A logic program, or knowledge base, consists of a collection of facts \(F~\\widetilde\{F\}\) and a collection of rules \(Q~\\widetilde\{Q\}\) that define the system’s knowledge\. For each test case, MetaSpace constructs a logic program combining observed facts with MR rules:

\(1\)\(𝑃𝑟𝑜𝑔𝑟𝑎𝑚\)\\displaystyle\\mathit\{\(Program\)\}𝒫spatial\\displaystyle\\quad\\mathcal\{P\}\_\{\\text\{spatial\}\}::=\\displaystyle\{:=\}F~observations\+\+Q~MRs\\displaystyle\\quad\\widetilde\{F\}\_\{\\text\{observations\}\}\\,\{\+\}\{\+\}\\,\\widetilde\{Q\}\_\{\\text\{MRs\}\}Here, the tilde notation indicates a list of items\. The Prolog engine then validates consistency by checking whether agent responses satisfy the expected spatial relations derived from the rules\. Any inconsistency indicates a spatial cognition violation\.

## 3\.Metamorphic Testing

### 3\.1\.Motivation for Using MT

A fundamental challenge in evaluating embodied agents is thetest oracle problem, i\.e\., the difficulty of determining the correct, expected output for a given test case\. Embodied tasks occur in dynamic environments, requiring real\-time route planning as environmental conditions evolve\. Therefore, evaluating the correctness and optimality of answers generated by embodied agents presents a significant challenge\. Existing research often assesses embodied intelligence through MCQ benchmarks\. These benchmarks typically consist of a set of images accompanied by manually designed questions, usually derived from human\-generated scenarios\. During the evaluation process, the agent is presented with the image, the corresponding question, and a selection of multiple\-choice answers\. However, the consistency and the quality of benchmark questions can vary significantly due to the differing levels of expertise among the human experts involved in their creation, which is not only labor\-intensive but also prone to inconsistencies\. Moreover, human\-generated test cases may not adequately cover edge cases, as it is inherently challenging for human experts to annotate answers for tasks involving constantly changing scenarios\.

MT offers a powerful solution to the test oracle problem and the limitations of manual test generation in dynamic environments\. Inspired by its significant success in assessing the quality of software\(Mansur et al\.,[2021](https://arxiv.org/html/2608.07533#bib.bib35); Tolksdorf et al\.,[2019](https://arxiv.org/html/2608.07533#bib.bib57)\), AI models\(Wang and Su,[2020](https://arxiv.org/html/2608.07533#bib.bib61)\), compilers\(Xiao et al\.,[2022](https://arxiv.org/html/2608.07533#bib.bib67),[2025](https://arxiv.org/html/2608.07533#bib.bib66)\), and quantum computing platforms\(Paltenghi and Pradel,[2023](https://arxiv.org/html/2608.07533#bib.bib45)\), we adapt MT to evaluate the embodied spatial cognition of embodied intelligence\. The core strength of MT is its ability to validate outputs without a predefined oracle\. Instead of asserting the exact output, it uses MRs to check for expected consistencies\. For instance, to test the implementation ofsin⁡\(x\)\\sin\(x\), we do not need to know the expected output for arbitrary floating\-point inputsxx\. Instead, we can assert that theMR:sin⁡\(x\)=sin⁡\(π−x\)\\text\{MR\}:\\sin\(x\)=\\sin\(\\pi\-x\)must always hold\. A discrepancy between the outputs forsin⁡\(x\)\\sin\(x\)andsin⁡\(π−x\)\\sin\(\\pi\-x\)reveals an implementation error, effectively bypassing the oracle problem\.

### 3\.2\.Formulation of Conducting MT with MetaSpace

We formalize the MT process in MetaSpace as follows\. In MT, a MR defines a predictable relationship between the outputs of an embodied agent when its inputs are mutated\. Specifically, for a functionf:ℐ→𝒪f:\\mathcal\{I\}\\to\\mathcal\{O\}, whereℐ\\mathcal\{I\}denotes the inputs \(e\.g\., visual observations\),𝒪\\mathcal\{O\}represents the outputs \(e\.g\., perceived spatial relations, movement estimations\), andffis the embodied agent under test\. Based on a given MR, we perform a transformationϕ:ℐ→ℐ\\phi:\\mathcal\{I\}\\to\\mathcal\{I\}on an original input𝐱∈ℐ\\mathbf\{x\}\\in\\mathcal\{I\}to generate a new input𝐱′=ϕ​\(𝐱\)∈ℐ\\mathbf\{x\}^\{\\prime\}=\\phi\(\\mathbf\{x\}\)\\in\\mathcal\{I\}\. The corresponding outputs are𝐲=f​\(𝐱\)\\mathbf\{y\}=f\(\\mathbf\{x\}\)and𝐲′=f​\(𝐱′\)\\mathbf\{y\}^\{\\prime\}=f\(\\mathbf\{x\}^\{\\prime\}\)\.

The MR is encoded as a boolean predicateMR​\(𝐱,𝐱′,𝐲,𝐲′\)\\mathrm\{MR\}\(\\mathbf\{x\},\\mathbf\{x\}^\{\\prime\},\\mathbf\{y\},\\mathbf\{y\}^\{\\prime\}\), which evaluates toTrue\\mathrm\{True\}if and only if𝐲\\mathbf\{y\}and𝐲′\\mathbf\{y\}^\{\\prime\}satisfy a predefined consistency constraint𝒫ϕ​\(𝐲,𝐲′\)\\mathcal\{P\}\_\{\\phi\}\(\\mathbf\{y\},\\mathbf\{y\}^\{\\prime\}\)based on principled logical rules and physical laws:

\(2\)MR​\(𝐱,𝐱′,𝐲,𝐲′\)=\{True⇔𝒫ϕ​\(𝐲,𝐲′\),Falseotherwise\.\\mathrm\{MR\}\(\\mathbf\{x\},\\mathbf\{x\}^\{\\prime\},\\mathbf\{y\},\\mathbf\{y\}^\{\\prime\}\)=\\begin\{cases\}\\mathrm\{True\}&\\iff\\mathcal\{P\}\_\{\\phi\}\(\\mathbf\{y\},\\mathbf\{y\}^\{\\prime\}\),\\\\ \\mathrm\{False\}&\\text\{otherwise\}\.\\end\{cases\}Overall, the general process of conducting MT using MetaSpace involves the following steps:

1. \(1\)Select an initial input𝐱∈ℐ\\mathbf\{x\}\\in\\mathcal\{I\}and obtain the output𝐲=f​\(𝐱\)\\mathbf\{y\}=f\(\\mathbf\{x\}\)\.
2. \(2\)Apply a metamorphic transformationϕ\\phicompatible with𝐱\\mathbf\{x\}to generate𝐱′=ϕ​\(𝐱\)∈ℐ\\mathbf\{x\}^\{\\prime\}=\\phi\(\\mathbf\{x\}\)\\in\\mathcal\{I\}\), and obtain output𝐲′=f​\(𝐱′\)\\mathbf\{y\}^\{\\prime\}=f\(\\mathbf\{x\}^\{\\prime\}\)\.
3. \(3\)Check whether the tuple\(𝐱,𝐱′,𝐲,𝐲′\)\(\\mathbf\{x\},\\mathbf\{x\}^\{\\prime\},\\mathbf\{y\},\\mathbf\{y\}^\{\\prime\}\)satisfies the predefined MR\. If not, it indicates a potential error in the agent’s spatial cognitive responses\.

### 3\.3\.MRs in MetaSpace

MetaSpace implements MRs derived from principled logic rules and physical laws to find potential errors in the outputs of embodied intelligence models\. We employ logic programming to implement these MRs \(discussed soon in[Section4\.3](https://arxiv.org/html/2608.07533#S4.SS3)\)\. These MRs holistically capture diverse spatial cognitive capabilities in embodied scenarios mentioned in[Fig\.2](https://arxiv.org/html/2608.07533#S2.F2)\. Additionally, as we will demonstrate in §[6\.4\.1](https://arxiv.org/html/2608.07533#S6.SS4.SSS1), these MRs are highly accurate in identifying spatial cognition errors in embodied agents\.

Table 1\.MRs for Spatial Cognition Evaluation \(Please refer to Fig\.[2](https://arxiv.org/html/2608.07533#S2.F2)for the spatial cognitive capabilities associated with the codes in the table\)\.Our MRs can be broadly categorized into two groups based on their foundational principles: logical consistency\-oriented MRs \([Section3\.3\.1](https://arxiv.org/html/2608.07533#S3.SS3.SSS1)\) and physical law\-oriented MRs \([Section3\.3\.2](https://arxiv.org/html/2608.07533#S3.SS3.SSS2)\)\. Table[1](https://arxiv.org/html/2608.07533#S3.T1)summarizes the six MRs implemented in MetaSpace, along with their foundational principles and the specific spatial cognitive capabilities they assess \(as defined in Fig\.[2](https://arxiv.org/html/2608.07533#S2.F2)\)\. In[Section3\.3\.1](https://arxiv.org/html/2608.07533#S3.SS3.SSS1)and[Section3\.3\.2](https://arxiv.org/html/2608.07533#S3.SS3.SSS2), we provide a detailed introduction to the MRs designed in MetaSpace\.

#### 3\.3\.1\.Logical Consistency\-Oriented MRs

MRs based on logical consistency are designed to ensure that the outputs of original and transformed test cases adhere to predefined logical rules\. These MRs are grounded in fundamental principles of logic, and they help identify inconsistencies in the spatial cognitive responses of embodied agents\.

- •MR1: Transitivity\.This MR is grounded in the transitivity property of directional spatial relations\. It asserts that any violation of the transitive property defined by Definition[3\.1](https://arxiv.org/html/2608.07533#S3.Thmtheorem1)indicates a potential spatial cognition error regarding directional relationships\. ###### Definition 0 \(MR1: Transitivity\)\. Lets1,s2,s3s\_\{1\},s\_\{2\},s\_\{3\}be three states\. In MetaSpace,sis\_\{i\}can represent either \(i\) the spatial state of an object, whereD​i​r​\(si,sj\)Dir\(s\_\{i\},s\_\{j\}\)denotes the spatial \(e\.g\., directional\) relation between objects’ statessis\_\{i\}andsjs\_\{j\}\(e\.g\., north, east, southwest, above\), or \(ii\) the state of an agent, whereD​i​r​\(si,sj\)Dir\(s\_\{i\},s\_\{j\}\)denotes the agent’s perceived movement direction from statesis\_\{i\}tosjs\_\{j\}\(e\.g\., go forward, backward, upward\)\. The operator⊕\\oplusdenotes the composition of such relations\. Given three state pairs\(s1,s2\)\(s\_\{1\},s\_\{2\}\),\(s2,s3\)\(s\_\{2\},s\_\{3\}\), and\(s1,s3\)\(s\_\{1\},s\_\{3\}\), the relations among these pairs must satisfy the transitive property, defined as: \(3\)MR1\(s1,s2,s3\):Dir\(s1,s3\)=Dir\(s1,s2\)⊕Dir\(s2,s3\)MR\_\{1\}\(s\_\{1\},s\_\{2\},s\_\{3\}\):\\quad Dir\(s\_\{1\},s\_\{3\}\)=Dir\(s\_\{1\},s\_\{2\}\)\\oplus Dir\(s\_\{2\},s\_\{3\}\)That is, the direction or motion froms1s\_\{1\}tos3s\_\{3\}must equal the composition of the one froms1s\_\{1\}tos2s\_\{2\}and froms2s\_\{2\}tos3s\_\{3\}, ensuring logical consistency\. Examples illustrating MR1 are shown in Fig\.[3\(a\)](https://arxiv.org/html/2608.07533#S3.F3.sf1)\.Example 1:With this MR, forssrepresenting an object’s state, if the directional relation betweens1s\_\{1\}ands2s\_\{2\}is east, and betweens2s\_\{2\}ands3s\_\{3\}is north, then the relation betweens1s\_\{1\}ands3s\_\{3\}must be northeast \(i\.e\., east⊕\\oplusnorth==northeast\)\.Example 2:Similarly, if the directional relation betweens1s\_\{1\}ands2s\_\{2\}is northeast, and betweens2s\_\{2\}ands3s\_\{3\}is north, then the relation betweens1s\_\{1\}ands3s\_\{3\}must still be northeast \(i\.e\., northeast⊕\\oplusnorth==northeast\)\.Example 3:Moreover, forssrepresenting an agent’s state, if an agent perceives the movement froms1s\_\{1\}tos2s\_\{2\}as moving forward, and froms2s\_\{2\}tos3s\_\{3\}as moving right, then the agent’s response to the movement froms1s\_\{1\}tos3s\_\{3\}must be equivalent to the combination of these two movements \(i\.e\., move forward⊕\\oplusmove right\)\. ![Refer to caption](https://arxiv.org/html/2608.07533v1/x5.png)\(a\)MR1: Transitivity ![Refer to caption](https://arxiv.org/html/2608.07533v1/x6.png)\(b\)MR2: Symmetry Figure 3\.Examples of MR1 and MR2 in both object spatial relations and agent movement perceptions\.
- •MR2: Symmetry\.MR2 leverages the symmetry property in spatial relations\. The directional relation or motion should satisfy the antisymmetry property, while the magnitude \(distance\) relation should satisfy the symmetry property, as defined in Definition[3\.2](https://arxiv.org/html/2608.07533#S3.Thmtheorem2)\. ###### Definition 0 \(MR2: Symmetry\)\. LetD​i​s​t​\(s1,s2\)Dist\(s\_\{1\},s\_\{2\}\)denote the distance froms1s\_\{1\}tos2s\_\{2\}, andI​n​vInvbe the inverse operator for directional relations or movement directions\. Given a state pair\(s1,s2\)\(s\_\{1\},s\_\{2\}\), the direction or motion froms2s\_\{2\}tos1s\_\{1\}must be the inverse of the one froms1s\_\{1\}tos2s\_\{2\}, and the distance froms2s\_\{2\}tos1s\_\{1\}must equal the distance froms1s\_\{1\}tos2s\_\{2\}, as defined: \(4\)MR2\(s1,s2\):\{D​i​r​\(s2,s1\)=I​n​v​\(D​i​r​\(s1,s2\)\)D​i​s​t​\(s2,s1\)=D​i​s​t​\(s1,s2\)MR\_\{2\}\(s\_\{1\},s\_\{2\}\):\\quad\\begin\{cases\}Dir\(s\_\{2\},s\_\{1\}\)=Inv\(Dir\(s\_\{1\},s\_\{2\}\)\)\\\\ Dist\(s\_\{2\},s\_\{1\}\)=Dist\(s\_\{1\},s\_\{2\}\)\\end\{cases\} We show examples illustrating MR2 in Fig\.[3\(b\)](https://arxiv.org/html/2608.07533#S3.F3.sf2)\.Example 1:If the direction betweens1s\_\{1\}ands2s\_\{2\}is east, then the direction betweens2s\_\{2\}ands1s\_\{1\}must be west \(i\.e\.,I​n​v​\(east\)=westInv\(\\text\{east\}\)=\\text\{west\}\)\.Example 2:Similarly, if the motion froms1s\_\{1\}tos2s\_\{2\}is moving forward, then the motion froms2s\_\{2\}tos1s\_\{1\}must be moving backward \(i\.e\.,I​n​v​\(move forward\)=move backwardInv\(\\text\{move forward\}\)=\\text\{move backward\}\)\.Example 3:If the distance betweens1s\_\{1\}ands2s\_\{2\}is 3 steps, then the distance betweens2s\_\{2\}ands1s\_\{1\}must also be 3 steps \(i\.e\.,D​i​s​t​\(s2,s1\)=D​i​s​t​\(s1,s2\)=3Dist\(s\_\{2\},s\_\{1\}\)=Dist\(s\_\{1\},s\_\{2\}\)=3steps\)\.444MetaSpace employs steps as distance units, as they are more relevant to embodied contexts\.
- •MR3: Contradiction\.MR3 is based on the law of non\-contradiction, which states that contradictory statements cannot be true simultaneously, as defined in Definition[3\.3](https://arxiv.org/html/2608.07533#S3.Thmtheorem3)\. It applies to both directional and magnitude spatial relationships\. ###### Definition 0 \(MR3: Contradiction\)\. It is impossible fors1s\_\{1\}to be in two different directions froms2s\_\{2\}at the same time, and similarly, the distance betweens1s\_\{1\}ands2s\_\{2\}cannot simultaneously be two different values\. The MR is formally defined as: \(5\)MR3\(s1,s2\):\{¬\(D​i​r​\(s1,s2\)=d1∧D​i​r​\(s1,s2\)=d2\),∀d1≠d2¬\(D​i​s​t​\(s1,s2\)=m1∧D​i​s​t​\(s1,s2\)=m2\),∀m1≠m2MR\_\{3\}\(s\_\{1\},s\_\{2\}\):\\quad\\begin\{cases\}\\neg\\left\(Dir\(s\_\{1\},s\_\{2\}\)=d\_\{1\}\\wedge Dir\(s\_\{1\},s\_\{2\}\)=d\_\{2\}\\right\),\\quad\\forall d\_\{1\}\\neq d\_\{2\}\\\\ \\neg\\left\(Dist\(s\_\{1\},s\_\{2\}\)=m\_\{1\}\\wedge Dist\(s\_\{1\},s\_\{2\}\)=m\_\{2\}\\right\),\\quad\\forall m\_\{1\}\\neq m\_\{2\}\\end\{cases\} Example 1:The agent cannot consider the motion froms1s\_\{1\}tos2s\_\{2\}as both moving forward and moving backward\. In other words, it cannot justify that both answers are correct \(i\.e\.,¬\(move forward∧move backward\)\\neg\(\\text\{move forward\}\\wedge\\text\{move backward\}\)\)\.Example 2:The agent cannot simultaneously justify that the direction froms1s\_\{1\}tos2s\_\{2\}is both east and west \(i\.e\.,¬\(east∧west\)\\neg\(\\text\{east\}\\wedge\\text\{west\}\)\)\.Example 3:Furthermore, the agent cannot perceive the distance froms1s\_\{1\}tos2s\_\{2\}as both 2 steps and 3 steps \(i\.e\.,¬\(2​steps∧3​steps\)\\neg\(2\\text\{ steps\}\\wedge 3\\text\{ steps\}\)\)\. We illustrate MR3 in Fig\.[4\(a\)](https://arxiv.org/html/2608.07533#S3.F4.sf1)\.

![Refer to caption](https://arxiv.org/html/2608.07533v1/x7.png)\(a\)MR3: Contradiction
![Refer to caption](https://arxiv.org/html/2608.07533v1/x8.png)\(b\)MR4: Triangle Inequality

Figure 4\.Examples of MR3: Contradiction and MR4: Triangle Inequality\.
#### 3\.3\.2\.Physical Law\-Oriented MRs

Physical law\-oriented MRs validate whether the spatial cognition remains consistent with physical laws\. They are defined as follows:

- •MR4: Triangle Inequality\.MR4 is based on the triangle inequality theorem\. The theorem states that in Euclidean space, for any triangle, the sum of the lengths of any two sides must be greater than or equal to the length of the remaining side\. Any violation of this theorem indicates a potential spatial cognition error by the agent, as defined in Definition[3\.4](https://arxiv.org/html/2608.07533#S3.Thmtheorem4)\. We illustrate MR4 in[Fig\.4\(b\)](https://arxiv.org/html/2608.07533#S3.F4.sf2)\. ###### Definition 0 \(MR4: Triangle Inequality\)\. Given three statess1,s2,s3s\_\{1\},s\_\{2\},s\_\{3\}, if they are collinear \(i\.e\., lie on the same straight line\), then whether considering the spatial distance between objects or the movement distance of an agent, the distance betweens1s\_\{1\}ands3s\_\{3\}equals either the sum or the absolute difference of the other two distances, depending on their relative directions\. Otherwise, the distance betweens1s\_\{1\}ands3s\_\{3\}must be strictly less than the sum of the other two distances\. \(6\)M​R4​\(s1,s2,s3\):\{D​i​s​t​\(s1,s3\)=D​i​s​t​\(s1,s2\)\+D​i​s​t​\(s2,s3\),if​D​i​r​\(s1,s2\)=D​i​r​\(s2,s3\)D​i​s​t​\(s1,s3\)=\|D​i​s​t​\(s1,s2\)−D​i​s​t​\(s2,s3\)\|,if​D​i​r​\(s1,s2\)=I​n​v​\(D​i​r​\(s2,s3\)\)D​i​s​t​\(s1,s3\)<D​i​s​t​\(s1,s2\)\+D​i​s​t​\(s2,s3\),otherwiseMR\_\{4\}\(s\_\{1\},s\_\{2\},s\_\{3\}\):\\begin\{cases\}Dist\(s\_\{1\},s\_\{3\}\)=Dist\(s\_\{1\},s\_\{2\}\)\+Dist\(s\_\{2\},s\_\{3\}\),&\\text\{if \}Dir\(s\_\{1\},s\_\{2\}\)=Dir\(s\_\{2\},s\_\{3\}\)\\\\ Dist\(s\_\{1\},s\_\{3\}\)=\\left\|Dist\(s\_\{1\},s\_\{2\}\)\-Dist\(s\_\{2\},s\_\{3\}\)\\right\|,&\\text\{if \}Dir\(s\_\{1\},s\_\{2\}\)=Inv\(Dir\(s\_\{2\},s\_\{3\}\)\)\\\\ Dist\(s\_\{1\},s\_\{3\}\)<Dist\(s\_\{1\},s\_\{2\}\)\+Dist\(s\_\{2\},s\_\{3\}\),&\\text\{otherwise\}\\end\{cases\}
- •MR5: Size\-Depth Consistency\.This relation is grounded in the principles of perspective geometry and is used to validate the agent’s perception of depth\. ###### Definition 0 \(MR5: Size\-Depth Consistency\)\. Letoobe a specific object observed by the agent in two different frames,f1f\_\{1\}andf2f\_\{2\}\. LetS​\(fi,o\)S\(f\_\{i\},o\)denote the projected size \(e\.g\., pixel height\) of objectooin framefif\_\{i\}, andD​\(fi,o\)D\(f\_\{i\},o\)denote the depth \(distance from the agent\) ofooin framefif\_\{i\}\. Under the assumption of a pinhole camera model, perspective geometry givesS​\(f,o\)∝1D​\(f,o\)S\(f,o\)\\propto\\frac\{1\}\{D\(f,o\)\}\. Thus, for two frames: \(7\)MR5\(f1,f2,o\):\|S​\(f1,o\)S​\(f2,o\)−D​\(f2,o\)D​\(f1,o\)\|<ϵMR\_\{5\}\(f\_\{1\},f\_\{2\},o\):\\quad\\left\|\\frac\{S\(f\_\{1\},o\)\}\{S\(f\_\{2\},o\)\}\-\\frac\{D\(f\_\{2\},o\)\}\{D\(f\_\{1\},o\)\}\\right\|<\\epsilonThat is, the ratio of projected sizes should be inversely proportional to the ratio of depths, within a relative error toleranceϵ\\epsilon\. [Fig\.5\(a\)](https://arxiv.org/html/2608.07533#S3.F5.sf1)illustrates this concept\. In this MR, we adopt the ideal pinhole camera model with consistent intrinsic parameters \(e\.g\., focal length, principal point\) as a standard approximation\. Although deviations may arise due to factors such as lens distortion, this assumption remains reasonable\. This is because current MLLM\-driven embodied intelligence primarily relies on RGB images as input and focuses on coarse\-grained depth estimation during embodied tasks rather than precise depth estimation\(Yang et al\.,[2025](https://arxiv.org/html/2608.07533#bib.bib70); Li et al\.,[2024b](https://arxiv.org/html/2608.07533#bib.bib28)\)\. Therefore, we accept controllable deviations\. Unlike the previous MRs, MR5 uses meters as the estimation unit\. To accommodate acceptable deviations in depth estimation, we introduce a relative tolerance thresholdϵ\\epsilon\. We discuss the threshold determination and sensitivity analysis in §[5\.5](https://arxiv.org/html/2608.07533#S5.SS5)\. ![Refer to caption](https://arxiv.org/html/2608.07533v1/x9.png)\(a\)MR5 ![Refer to caption](https://arxiv.org/html/2608.07533v1/x10.png)\(b\)MR6 Figure 5\.Under the pinhole camera model, perspective geometry dictates that the ratio of projected sizes should be inversely proportional to the ratio of depths for a single object\. MR5 \(Size\-Distance Consistency for Multi\-Frame Single Object\) asserts that a single object’s size\-depth relationship remains consistent across different frames, while MR6 \(Size\-Distance Consistency for Multi\-Object Multiple Frames\) ensures that the size ratio between different objects remains consistent across frames\.
- •MR6: Object Size Ratio Consistency\.This relation is used to validate the agent’s perception of object size, as defined in Definition[3\.6](https://arxiv.org/html/2608.07533#S3.Thmtheorem6)\. ###### Definition 0 \(MR6: Object Size Ratio Consistency\)\. Letf1f\_\{1\}andf2f\_\{2\}be two frames from the agent’s trajectory, ando1o\_\{1\}ando2o\_\{2\}be two distinct objects observed in both frames\. DenoteH​\(fi,oj\)H\(f\_\{i\},o\_\{j\}\)as the estimated real\-world size of objectojo\_\{j\}in framefif\_\{i\}\. The following approximate consistency relation should hold: \(8\)MR6\(f1,f2,o1,o2\):\|H​\(f1,o1\)H​\(f1,o2\)−H​\(f2,o1\)H​\(f2,o2\)\|<δMR\_\{6\}\(f\_\{1\},f\_\{2\},o\_\{1\},o\_\{2\}\):\\quad\\left\|\\frac\{H\(f\_\{1\},o\_\{1\}\)\}\{H\(f\_\{1\},o\_\{2\}\)\}\-\\frac\{H\(f\_\{2\},o\_\{1\}\)\}\{H\(f\_\{2\},o\_\{2\}\)\}\\right\|<\\deltaThat is, the ratio of the estimated real\-world sizes of two objects should remain consistent across different frames within a relative error toleranceδ\\delta\. We discuss the determination and sensitivity analysis of this threshold in §[5\.5](https://arxiv.org/html/2608.07533#S5.SS5)\. The reason for comparing a ratio instead of absolute size is that we observe the scale of the virtual simulator environment is often altered by the stretching or scaling of 3D models\.

### 3\.4\.Discussions on MR Design

Our set of six MRs is not an arbitrary collection of heuristics; rather, it is derived through a systematic selection process governed by three rigorous criteria: \(1\)Theoretical Foundation\.Each MR must be grounded in eitheruniversal logical axiomsorfundamental physical laws\. This ensures MRs are model\-agnostic and unbiased toward specific architectures\. \(2\)Testability\.Each MR must be verifiable using only spatiotemporal state data available during embodied execution\. This excludes purely cognitive phenomena that lack observable behavioral correlates\. \(3\)Coverage of Core Spatial Attributes\.The MR set must collectively test the four fundamental spatial attributes \(direction, distance, size, depth\) identified in §[2\.1](https://arxiv.org/html/2608.07533#S2.SS1)\.

Based on these criteria, we derived aminimal sufficient setof MRs to cover the embodied SC spectrum\. Our set fully covers the four fundamental SCs identified in §[2\.1](https://arxiv.org/html/2608.07533#S2.SS1)without redundancy to ensure the completeness of our MR set, as each MR targets distinct cognitive mechanisms\. For instance, within directional tasks, MR1 \(Transitivity\) focuses on multi\-hop reasoning \(e\.g\.,s1→s2→s3s\_\{1\}\\to s\_\{2\}\\to s\_\{3\}\), whereas MR2 \(Symmetry\) assesses single\-step reasoning\. Similarly, within magnitude tasks, MR6 \(Object Size Ratio Consistency\) evaluates relative perception between objects’ sizes, whereas MR5 \(Size\-Depth Consistency\) demands adherence to perspective geometry linking projected size to depth\. Regarding reliability, as demonstrated in[Section6\.2](https://arxiv.org/html/2608.07533#S6.SS2), the high human baseline \(0\.96\) confirms that our MRs accommodate valid spatial interpretations, ensuring minimal over\-flagging\. We further provide a quantitative false positive analysis in[Section6\.4\.2](https://arxiv.org/html/2608.07533#S6.SS4.SSS2)\. Finally, our MR design ensures universality\. By grounding MRs in fundamental spatial attributes rather than specific tasks, they remain applicable across diverse SC types\. For instance, MR1 \(Transitivity\) applies equally to object relations \(SC2\-a\) and motion perception \(SC1\-a\), ensuring adaptability even as new SC tasks emerge\. We empirically validate the effectiveness of these MRs in[Section6\.4\.1](https://arxiv.org/html/2608.07533#S6.SS4.SSS1)\.

![Refer to caption](https://arxiv.org/html/2608.07533v1/x11.png)Figure 6\.Pipeline of the MetaSpace framework for evaluating embodied spatial cognition\.

## 4\.Methodology

We design and implement MetaSpace to address the challenges mentioned above\. The overall architecture of MetaSpace consists of four main modules, as illustrated in[Fig\.6](https://arxiv.org/html/2608.07533#S3.F6):

- •Spatiotemporal State Data Collection \([Section4\.1](https://arxiv.org/html/2608.07533#S4.SS1)\):Collects and preprocesses real\-world trajectory data from embodied agents, sampling state combinations and extracting objects for downstream test case generation\.
- •Test Case Generation \([Section4\.2](https://arxiv.org/html/2608.07533#S4.SS2)\):Automatically generates original and transformed test cases from state combinations using metamorphic transformations based on predefined MRs\.
- •Automated Validation via Logic Programming \([Section4\.3](https://arxiv.org/html/2608.07533#S4.SS3)\):Encodes agent responses and MRs as Prolog facts and rules, and validates consistency between original and transformed outputs to detect violations in embodied spatial cognition responses\.
- •Spatial Capability Scoring and Analysis \([Section4\.4](https://arxiv.org/html/2608.07533#S4.SS4)\):Aggregates violation statistics and computes scores for each spatial cognition capability, enabling detailed evaluation and comparison\.

### 4\.1\.Spatiotemporal State Data Collection

In MetaSpace, we utilize real\-world trajectories of embodied intelligence to evaluate spatial cognition capabilities, ensuring a strong correlation between evaluation results and actual embodied task performance\. Therefore, instead of relying on static, manually curated test cases, our framework processes real execution trajectories from embodied agents \(more details in[Section5\.1](https://arxiv.org/html/2608.07533#S5.SS1)\)\. Let𝒯\\mathcal\{T\}represent the set of raw trajectory data collected from embodied agents performing embodied tasks \(e\.g\., navigation or manipulation\), defined as𝒯=\{𝒯1,𝒯2,…,𝒯m\}\\mathcal\{T\}=\\\{\\mathcal\{T\}\_\{1\},\\mathcal\{T\}\_\{2\},\\ldots,\\mathcal\{T\}\_\{m\}\\\}, where each trajectory𝒯k\\mathcal\{T\}\_\{k\}is a sequence of states defined as𝒯k=\{sk,1,sk,2,…,sk,n\}\\mathcal\{T\}\_\{k\}=\\\{s\_\{k,1\},s\_\{k,2\},\\ldots,s\_\{k,n\}\\\}\. Each statesk,is\_\{k,i\}includes an egocentric visual observation by the embodied agent, represented asfk,if\_\{k,i\}\(image frame\)\. These states are then used to generate test cases \(detailed in[Section4\.2](https://arxiv.org/html/2608.07533#S4.SS2)\)\. We preprocess the trajectory data for test case generation in the following:

➀ Trajectory Sampling\.Each trajectory𝒯k\\mathcal\{T\}\_\{k\}is processed to extract state combinations \(e\.g\., two or three states\) based on the specific SC and MR utilized\. Each combination is then used for individual test case generation\. For instance, to adaptMR2: SymmetryinSC1, we extract pairs of consecutive states\(sk,i,sk,i\+1\)\(s\_\{k,i\},s\_\{k,i\+1\}\)from the trajectory𝒯k\\mathcal\{T\}\_\{k\}\. The implementation details of state combination extraction are provided in[Section5\.3](https://arxiv.org/html/2608.07533#S5.SS3)\.

➁ Object Detection and Tracking\.For each state combination, we extract the objects \(e\.g\., landmarks, manipulable items\) for test case generation\. We implement this using Ultralytics YOLO\(Redmon et al\.,[2016](https://arxiv.org/html/2608.07533#bib.bib49)\)for multi\-object tracking across the whole trajectory\. Let𝒪k,i=\{ok,i,1,ok,i,2,…,ok,i,l\}\\mathcal\{O\}\_\{k,i\}=\\\{o\_\{k,i,1\},o\_\{k,i,2\},\\ldots,o\_\{k,i,l\}\\\}denote the set of objects detected in a framefk,if\_\{k,i\}within trajectory𝒯k\\mathcal\{T\}\_\{k\}\. Each objectok,i,jo\_\{k,i,j\}is detected with its bounding box coordinates, class label, and confidence score\. Since the objective of MetaSpace is to evaluate embodied spatial cognition capabilities, we aim to minimize interference from perception capability failures \(e\.g\., the agent failing to recognize the target apple\)\. Therefore, we adopt the following strategies: 1\) Assign a unique identifier \(ID\) to each object across frames\. 2\) Retain only detections with a confidence score above 0\.6 to minimize noise \(the threshold determination process is detailed in[Section5\.5](https://arxiv.org/html/2608.07533#S5.SS5)\)\. 3\) Incorporate the detected bounding boxes into the input prompt for the embodied agent to further reduce the impact of perception errors\. We further investigate the performance of YOLO and the robustness against perception noise in[Section5\.3](https://arxiv.org/html/2608.07533#S5.SS3)\.

### 4\.2\.Test Case Generation

The state combinations and their corresponding object sets obtained from the[Section4\.1](https://arxiv.org/html/2608.07533#S4.SS1)are used to generate test cases\. In this module, the trajectories are utilized to generate test cases based on the MRs defined in[Section3\.3\.1](https://arxiv.org/html/2608.07533#S3.SS3.SSS1)and[Section3\.3\.2](https://arxiv.org/html/2608.07533#S3.SS3.SSS2)\. Each test case consists of an original input queryxxand a transformed input queryx′=ϕ​\(x\)x^\{\\prime\}=\\phi\(x\), which is generated by applying the metamorphic transformationϕ\\phiassociated with a specific type of MR\. Both are derived from the same state combination\.

Algorithm 1Automated Test Case Generation1:Trajectory set

𝒯\\mathcal\{T\}, Prompt Templates

t​e​m​p​l​a​t​e~\\widetilde\{template\}, target spatial cognition

S​CSC
2:Test Cases

𝒞\\mathcal\{C\}
3:

𝒞←\[\]\\mathcal\{C\}\\leftarrow\[\\,\],

M​R~←\\widetilde\{MR\}\\leftarrowGetMRsBySC\(

S​CSC\)

4:foreach trajectory

𝒯k\\mathcal\{T\}\_\{k\}in

𝒯\\mathcal\{T\}do

5:foreach

M​RMRin

M​R~\\widetilde\{MR\}do

6:if

S​C=SC1SC=\\text\{SC1\}then

7:

w←w\\leftarrowGetWindowSize\(

M​RMR\)

8:

𝒮←\\mathcal\{S\}\\leftarrowGenerateConsecutiveCombinations\(

𝒯k\\mathcal\{T\}\_\{k\},

ww\)

9:else

10:

𝒫←\\mathcal\{P\}\\leftarrowGetPersistentObjects\(

𝒯k\\mathcal\{T\}\_\{k\}\)

11:

𝒮←\\mathcal\{S\}\\leftarrowGenerateObjectBasedCombinations\(

𝒯k,𝒫\\mathcal\{T\}\_\{k\},\\mathcal\{P\}\)

12:endif

13:foreach state combination

SSin

𝒮\\mathcal\{S\}do

14:

f​i​l​l​e​d​\_​p​r​o​m​p​t←filled\\\_prompt\\leftarrowFillTemplate\(

t​e​m​p​l​a​t​e~​\[S​C,M​R\]\\widetilde\{template\}\[SC,MR\],

SS,

S​CSC\)

15:

\(x,x′\)←\(x,x^\{\\prime\}\)\\leftarrowGenerateQueries\(

f​i​l​l​e​d​\_​p​r​o​m​p​tfilled\\\_prompt,

M​RMR\)

16:

𝒞\.append​\(\(x,x′,M​R\)\)\\mathcal\{C\}\.\\text\{append\}\(\(x,x^\{\\prime\},MR\)\)
17:endfor

18:endfor

19:endfor

20:return

𝒞\\mathcal\{C\}

MetaSpace adopts an automated and spatial cognitive capability\-aware approach to sample trajectories and generate test cases\. The detailed process is outlined in Algorithm[1](https://arxiv.org/html/2608.07533#alg1)\. For each trajectory \(𝒯k\\mathcal\{T\}\_\{k\}\) in the trajectory set \(𝒯\\mathcal\{T\}\), and for each MR associated with the target SC, the algorithm selects state combinations using different strategies based on the SC type\. Specifically, for movement perception \(SC1\), a sliding window of sizeww\(determined by the MR\) is applied to extract consecutive state combinations \(Lines 5–6\)\. For SC2, SC3, and SC4, the algorithm first identifies persistent objects across the trajectory and then generates object\-centric state combinations based on these objects \(Lines 8–9\)\. For each state combination, the appropriate prompt template is retrieved from the predefined sett​e​m​p​l​a​t​e~​\[S​C,M​R\]\\widetilde\{template\}\[SC,MR\]and filled with the state combinations and objects \(objects are excluded for SC1\) \(Lines 12\)\. Subsequently, the original input queryxxis generated from the filled prompt \(Line 13\)\. The transformed input queryx′x^\{\\prime\}is created by applying the metamorphic transformationϕ\\phiassociated with the current MR to the filled prompt \(Line 13\)\. Finally, the pair\(x,x′\)\(x,x^\{\\prime\}\), along with the corresponding MR, is appended to the test case set𝒞\\mathcal\{C\}\(Line 14\)\. After processing all trajectories and MRs, the algorithm returns the complete set of generated test cases𝒞\\mathcal\{C\}\(Line 18\)\. This approach ensures that test case generation is fully automated, reproducible, and scalable, while remaining agnostic to the specific MR or the length of the agent’s trajectories\.

### 4\.3\.Automated Validation via Logic Programming

MetaSpace automates output validation by encoding spatial observations as Prolog facts and MRs as rules\. For each test case, the Prolog engine infers the expected output for the transformed input using these facts and rules, then compares it to the agent’s actual response\. Any inconsistency indicates a spatial cognition error\.

As detailed in[Algorithm2](https://arxiv.org/html/2608.07533#alg2), the process begins with an automatic rule parser that iterates over all MRs, extracting for each relation a specific query pattern \(a Prolog predicate template for enumerating relevant test case instances\) and the corresponding reasoning rule \(line 3\)\. A Prolog program is constructed by combining the facts from observations with the reasoning rule for the MR \(line 4\)\. Using this program, all possible instantiations of the MR predicate template, which represent valid combinations of states or objects present in the ground facts, are enumerated for validation \(line 5\)\. For each instantiation, the algorithm checks consistency between the expected output inferred by the Prolog engine and the agent’s actual response; if the check fails, a violation is recorded and appended to a set for further analysis \(lines 6–9\)\. This automated process enables MetaSpace to comprehensively assess the spatial cognitive consistency of the agent across all possible test cases derived from its observations\. We utilize SWI\-Prolog\(Wielemaker et al\.,[2012](https://arxiv.org/html/2608.07533#bib.bib63)\), an open\-source advanced logic programming interpreter to implement the Prolog engine\.

Algorithm 2Automated Consistency Validation via Logic Programming1:Observations

F~o​b​s\\widetilde\{F\}\_\{obs\}, Metamorphic Relations

M​R~\\widetilde\{MR\}
2:Violation Set

V~\\widetilde\{V\}
3:

V~←\[\]\\widetilde\{V\}\\leftarrow\[\\,\]⊳\\trianglerightInitialization

4:foreach

M​RMRin

M​R~\\widetilde\{MR\}do⊳\\trianglerightIterate over each MR

5:

\(QM​R,ℛM​R\)←\(Q\_\{MR\},\\mathcal\{R\}\_\{MR\}\)\\leftarrowParseMR\(

M​RMR\)⊳\\trianglerightObtain MR\-specific query and reasoning rule

6:

𝒫←F~o​b​s\+\+ℛM​R\\mathcal\{P\}\\leftarrow\\widetilde\{F\}\_\{obs\}\+\+\\mathcal\{R\}\_\{MR\}⊳\\trianglerightConstruct Prolog program with facts and MR rule

7:

I​n​s​t←Inst\\leftarrowFindAllInstantiations\(

𝒫\\mathcal\{P\},

QM​RQ\_\{MR\}\)⊳\\trianglerightEnumerate all entity tuples to check

8:foreach

i​n​s​tinstin

I​n​s​tInstdo⊳\\trianglerightIterate over each instantiation

9:ifnotCheckConsistency\(

𝒫\\mathcal\{P\},

QM​RQ\_\{MR\},

i​n​s​tinst\)then

10:

Vn​e​w←\(M​R,i​n​s​t\)V\_\{new\}\\leftarrow\(MR,inst\)⊳\\trianglerightRecord violation: MR and instance

11:

V~\.append​\(Vn​e​w\)\\widetilde\{V\}\.\\text\{append\}\(V\_\{new\}\)
12:endif

13:endfor

14:endfor

15:return

V~\\widetilde\{V\}⊳\\trianglerightReturn all detected violations

### 4\.4\.Spatial Capability Scoring and Accuracy Analysis

After executing all test cases and collecting the violation setV~\\widetilde\{V\}, MetaSpace computes a spatial capability score for each spatial cognitive capability defined in Fig\.[2](https://arxiv.org/html/2608.07533#S2.F2)\. The scoring process is defined in the following\. For each embodied spatial cognitive capabilityS​CiSC\_\{i\}, identify the set of associated metamorphic relationsM​R~​\(S​Ci\)\\widetilde\{MR\}\(SC\_\{i\}\)\. Let𝒞~​\(S​Ci\)\{\\mathcal\{\\widetilde\{C\}\}\(SC\_\{i\}\)\}denote the set of all test cases generated forS​CiSC\_\{i\}based onM​R~​\(S​Ci\)\\widetilde\{MR\}\(SC\_\{i\}\), andV~​\(S​Ci\)\\widetilde\{V\}\(SC\_\{i\}\)denote the set of violations detected forS​CiSC\_\{i\}\. Compute the capability score using the formula:

\(9\)S​c​o​r​e​\(S​Ci\)=1−\|V~​\(S​Ci\)\|\|𝒞~​\(S​Ci\)\|Score\(SC\_\{i\}\)=1\-\\frac\{\|\\widetilde\{V\}\(SC\_\{i\}\)\|\}\{\|\\widetilde\{\\mathcal\{C\}\}\(SC\_\{i\}\)\|\}where\|𝒞~​\(S​Ci\)\|\|\\widetilde\{\\mathcal\{C\}\}\(SC\_\{i\}\)\|is the total number of test cases and\|V~​\(S​Ci\)\|\|\\widetilde\{V\}\(SC\_\{i\}\)\|is the number of violations\. This score ranges from 0 to 1, where a score of 1 indicates perfect performance \(no violations\), and a score of 0 indicates complete failure \(all test cases resulted in violations\)\.

## 5\.Implementation

### 5\.1\.Dataset

The three embodied scenarios and their corresponding datasets used in our evaluation are as follows:Household robot:We use the EB\-Navigation dataset from\(Yang et al\.,[2025](https://arxiv.org/html/2608.07533#bib.bib70)\), which is based on AI2\-THOR\(Kolve et al\.,[2017](https://arxiv.org/html/2608.07533#bib.bib25)\)\. It contains 60 navigation trajectories in different household scenes \(e\.g\., kitchens, living rooms, and bedrooms\)\.Robotic arm:We use the EB\-Manipulation dataset from\(Yang et al\.,[2025](https://arxiv.org/html/2608.07533#bib.bib70)\), which is based on VLMBench\(Zheng et al\.,[2022](https://arxiv.org/html/2608.07533#bib.bib80)\)using the CoppeliaSim simulator\(Rohmer et al\.,[2013](https://arxiv.org/html/2608.07533#bib.bib50)\)to control a 7\-DoF Franka Emika Panda robotic arm\. The dataset contains 48 manipulation trajectories in different tabletop scenes\.Drone:We use a sub\-dataset of drone navigation trajectories from\(Zhao et al\.,[2025](https://arxiv.org/html/2608.07533#bib.bib79)\), which is based on the AerialVLN\(Liu et al\.,[2023b](https://arxiv.org/html/2608.07533#bib.bib32)\)\. The sub\-dataset contains 50 navigation trajectories in various outdoor scenes \(e\.g\., urban areas, parks, and highways\)\.

### 5\.2\.Benchmark Agents

Model Selection\.To ensure reliable evaluation, we assess embodied agents powered by 6 SOTA MLLMs using MetaSpace\. We select two categories for analysis: \(1\) closed\-source, API\-accessible models, including GPT\-5\(OpenAI,[2025](https://arxiv.org/html/2608.07533#bib.bib43)\), GPT\-4o\(OpenAI,[2024](https://arxiv.org/html/2608.07533#bib.bib42)\), and Claude Sonnet 4\(Anthropic,[2025](https://arxiv.org/html/2608.07533#bib.bib3)\); and \(2\) open\-source, locally deployable models, including Qwen\-VL\(Wang et al\.,[2024](https://arxiv.org/html/2608.07533#bib.bib60)\), InternVL3\.5\-8B\(Wang et al\.,[2025](https://arxiv.org/html/2608.07533#bib.bib62)\), and DeepSeek\-VL2\-small\(Wu et al\.,[2024](https://arxiv.org/html/2608.07533#bib.bib65)\)\.

Model Configurations\.To ensure the stability and consistency of model outputs during evaluation, we set thetemperatureparameter to 0, resulting in deterministic responses\. We further settop\-pto 0\.9 and disabletop\-ksampling \(set to 0\), so that only the most probable tokens are selected, thereby improving the reliability of generated results\.

Consistency in Model Outputs\.To rigorously assess the consistency of MLLM responses in our approach, we conduct statistical significance tests\. Specifically, for each embodied scenario and SC, we randomly sample 30 test cases, yielding a total of 720 test cases\. Taking GPT\-4o as an example, each test case was executed five times under the above configuration\. To evaluate whether the responses from different runs were statistically indistinguishable, we then applied the Friedman test\(Friedman,[1937](https://arxiv.org/html/2608.07533#bib.bib19)\), a non\-parametric method for detecting differences across multiple repeated measures\. The results showed no significant differences between runs \(averagepp\-value = 0\.57\), confirming that the MLLM outputs are highly consistent under our settings\. Therefore, a single run is sufficient for evaluation in MetaSpace, ensuring both efficiency and reliability\.

Table 2\.MR5 Threshold Sensitivity Analysis \(Score\)\.
Table 3\.MR6 Threshold Sensitivity Analysis \(Score\)\.

### 5\.3\.Spatiotemporal State Data Collection

Trajectory Sampling Strategy\.The trajectory sampling strategy varies based on the SC capability being evaluated\. For movement perception evaluation, we exclusively utilize consecutive states to assess agents’ immediate motion perception rather than their long\-term path planning abilities\. This distinction enables embodied agents to support step\-by\-step action decisions through continuous perception\-action loops, while higher\-level planning capabilities rely on goal\-oriented strategies for long\-term navigation\. This approach emphasizes our focus on core SC skills rather than the extensively discussed path planning capabilities\(Kong et al\.,[2024](https://arxiv.org/html/2608.07533#bib.bib26); Zhang et al\.,[2024a](https://arxiv.org/html/2608.07533#bib.bib75)\)\. For evaluations of spatial reasoning, perspective visualization, and egocentric\-allocentric transformation, we do not restrict to consecutive states, as these capabilities have different requirements\. Spatial reasoning requires understanding the relationships between objects in space, often necessitating multiple time steps to observe the same objects from various angles\. Perspective visualization often requires observations from multiple viewpoints that may not be captured in consecutive frames\. Additionally, egocentric\-allocentric transformation requires agents to integrate information from different spatiotemporal positions to construct coherent mental maps\. Therefore, we extract different state combinations from the trajectories based on the SC being evaluated\.

Object Detection and Tracking\.We utilize YOLOv11\(Ultralytics,[2024](https://arxiv.org/html/2608.07533#bib.bib58)\)for multi\-object tracking in the collected spatiotemporal states\. Notably, MetaSpace utilizes YOLO solely for candidate object discovery to generate queries, rather than defining the ground\-truth of spatial relations\. The oracle of MetaSpace relies on internal logical consistency \(e\.g\., Transitivity\) or relative physical ratios, independent of absolute YOLO coordinates\. Therefore, potential missed objects \(due to recall limitations\) merely reduce the quantity of generated test cases without introducing false violations\. We evaluated YOLO’s tracking performance on a dataset of 500 images across 50 episodes\. The results demonstrate that YOLOv11 achieves 94% IoU without false positives, ensuring minimal perception errors\. To further validate robustness against perception noise, we injected synthetic noise \(including label inaccuracies and bounding box jitter of up to±20%\\pm 20\\%\) into 1,000 test instances\. We observed only six cognitive failures resulting from this noise, which were attributed to the MLLM’s inherent perception capabilities for self\-correction\. This confirms that perception noise has a negligible impact on our cognitive evaluation, and our modular design allows MetaSpace to adopt future perception advancements\. We discuss the threshold for YOLO in[Section5\.5](https://arxiv.org/html/2608.07533#S5.SS5)\.

### 5\.4\.Test Case Generation Implementation

In total, we generate 30,300 unique test cases\. This number is derived by exhaustively enumerating all valid state and object combinations across all scenarios, SC capabilities, and MRs, as determined by the automated test case generation process in[Section4\.2](https://arxiv.org/html/2608.07533#S4.SS2)\. These test cases comprehensively cover the eight SCs outlined in[Fig\.2](https://arxiv.org/html/2608.07533#S2.F2)\. For each SC, we create test cases based on the associated MRs defined in[Section3\.3\.1](https://arxiv.org/html/2608.07533#S3.SS3.SSS1)and[Section3\.3\.2](https://arxiv.org/html/2608.07533#S3.SS3.SSS2)\. We then utilize these test cases to evaluate the six benchmark embodied agents\.

### 5\.5\.Threshold Determination

To determine the optimal threshold for MR5 and MR6, the criterion is to ensure that the embodied agent achieves human\-level performance\. To establish a rigorous baseline, we recruited five adult participants, and each participant was presented with the same set of SC tasks and visual stimuli as the evaluated agents, under standardized instructions and conditions\. Participants completed the tasks independently, without time limits or external assistance, to minimize potential biases\. All responses were collected and evaluated using the same automated logic\-based validation framework applied to the agents\. The five participants exhibited high inter\-subject consistency with low standard deviations \(approx\.0\.050\.05for MR5 and approx\.0\.0250\.025for MR6\), indicating that increasing the participant count would unlikely shift the derived thresholds significantly\. By analyzing the relative errors between the participants’ estimates and the ground truth, the selected thresholds were calibrated to approximately the 90th percentile of human precision\. This allows for minor estimation deviations while rigorously capturing SC capabilities\. Consequently, we set the relative error threshold for MR5 \(ϵ\\epsilon\) and MR6 \(δ\\delta\) to 0\.1 and 0\.05, respectively\. This implies that we accept a 10% relative error for MR5 and 5% for MR6 as the bounds of human\-level performance; any estimates exceeding these thresholds are considered violations\. We further quantified robustness by testing human and GPT\-4o performance under threshold variations \(±25%\\pm 25\\%,±50%\\pm 50\\%\)\. As shown in Tab\.[3](https://arxiv.org/html/2608.07533#S5.T3)and Tab\.[3](https://arxiv.org/html/2608.07533#S5.T3), the GPT\-4o SC score improved by only approximately 0\.05 even when relaxing the threshold by 50%\. This suggests that the detected errors are catastrophic logic violations rather than marginal misses; therefore, slight changes in thresholds do not alter our research conclusions\. To determine the optimal confidence score threshold for multi\-object tracking using YOLO, we conduct experiments on 500 images from 50 episodes, testing thresholds ranging from 0\.4 to 0\.9\. Manual inspection of the detection results indicates that a threshold of 0\.6 yields an average of four detected objects per image, with a precision of 100% and a recall of 82%\. This threshold ensures both low false positive rates and adequate object coverage for spatial cognition evaluation\.

## 6\.Evaluation

Our evaluation aims to answer the following research questions \(RQs\):

- •RQ1 \(SC Performance\): How do different MLLM\-driven embodied agents perform in embodied spatial cognition?This question evaluates the embodied spatial cognition capabilities of various MLLM\-driven embodied agents by using MetaSpace\.
- •RQ2 \(Comparison with Existing Works\): How does MetaSpace compare with existing approaches in evaluating embodied SC capabilities?This RQ studies whether MetaSpace outperforms existing benchmarks from both quantitative and qualitative perspectives\.
- •RQ3 \(Internal Evaluation\): How effective are individual MRs and how reliable is the overall detection mechanism?This RQ investigates the effectiveness of each MR in detecting embodied spatial cognition errors and assesses the reliability of the entire detection framework\.
- •RQ4 \(Mitigation\): How can we mitigate the limitations of current MLLM\-driven embodied agents in spatial cognition tasks?This RQ explores potential strategies for improving the performance of embodied agents in spatial cognition tasks\.

### 6\.1\.Experimental Setup

Our experiments are conducted on a server running Ubuntu 22\.04, equipped with dual 56\-core Intel Xeon Scalable processors,2TB2\\text\{\\,\}\\mathrm\{T\}\\mathrm\{B\}of RAM, and an NVIDIA H800 GPU node\. The total GPU hours consumed for all experiments on open\-source MLLMs \(all scenarios\) amount to31 660\.412seconds31\\,660\.412\\text\{\\,\}\\mathrm\{s\}\\mathrm\{e\}\\mathrm\{c\}\\mathrm\{o\}\\mathrm\{n\}\\mathrm\{d\}\\mathrm\{s\}, which is acceptable given the scale of our evaluation\.

![Refer to caption](https://arxiv.org/html/2608.07533v1/x12.png)Figure 7\.Heatmap of embodied spatial cognition scores across different embodied agents\.
![Refer to caption](https://arxiv.org/html/2608.07533v1/x13.png)Figure 8\.Radar chart comparing embodied spatial cognition scores among different embodied agents\.

### 6\.2\.RQ1: SC Performance

To evaluate the embodied spatial cognition capabilities of six benchmark MLLM\-driven embodied agents, we analyze the statistics of test cases and violations detected by MetaSpace\. The detailed embodied spatial cognition scores for each agent across all SCs are displayed in Fig\.[8](https://arxiv.org/html/2608.07533#S6.F8)and Fig\.[8](https://arxiv.org/html/2608.07533#S6.F8)\. Additionally, Tab\.[4](https://arxiv.org/html/2608.07533#S6.T4)reports normalized error statistics\. The high failure density \(e\.g\.,\>\>82 errors per trajectory\) verifies that the substantial volume of detected errors stems from pervasive cognitive failures throughout task execution, rather than being an artifact of the test case scale\.

Magnitude vs\. Directional Tasks\.Results show that direction estimation and directional spatial reasoning tasks \(e\.g\., movement direction estimation in SC1\-a, directional reasoning in SC2\-a and SC4\-a\) are relatively challenging for most agents compared to magnitude estimation tasks \(e\.g\., distance estimation in SC1\-b, SC2\-b, SC4\-b, size estimation in SC3\-a, and depth estimation in SC3\-b\)\. This phenomenon can be attributed to the fact that magnitude\-related tasks primarily require direct numerical inference and quantitative reasoning based on visual input, without necessitating complex spatial relationship modeling\. Current mainstream MLLMs are pre\-trained on large\-scale corpora that emphasize static visual descriptions, object attribute recognition, and fundamental quantitative judgments, making these capabilities more readily generalizable to magnitude estimation tasks\(Chatterjee et al\.,[2024](https://arxiv.org/html/2608.07533#bib.bib8); Zhu et al\.,[2023](https://arxiv.org/html/2608.07533#bib.bib82)\)\. In contrast, direction estimation and directional spatial reasoning tasks require advanced spatial representations and reasoning chains in three\-dimensional space, which are not yet sufficiently developed in existing MLLMs\(Chatterjee et al\.,[2024](https://arxiv.org/html/2608.07533#bib.bib8); Hoehing et al\.,[2023](https://arxiv.org/html/2608.07533#bib.bib24); Ramakrishnan et al\.,[2025](https://arxiv.org/html/2608.07533#bib.bib48); Zhang et al\.,[2025c](https://arxiv.org/html/2608.07533#bib.bib73)\)\. As a result, while agents perform relatively well on magnitude tasks, their poor performance on directional SC highlights the limitations in spatial reasoning capabilities of current MLLM\-driven embodied agents\.

Size Estimation \(SC3\-a\)\.Among all spatial cognition tasks, most agents demonstrate outstanding performance in SC3\-a \(size estimation\), with all models except Qwen\-VL scoring above 0\.7\. This can be attributed to the robust spatial common sense knowledge base of MLLMs, which includes typical dimensions of common objects\. By referencing nearby objects with known sizes, MLLMs can effectively estimate the sizes of unknown objects\. This strategy renders size estimation tasks relatively easier for MLLM\-driven agents\.

Egocentric\-Allocentric Transformation \(SC4\)\.SC4 evaluates the agent’s ability to construct mental maps and perform reference frame transformations\. Similar to SC2 \(spatial reasoning\), SC4 involves reasoning about spatial relationships, but places greater emphasis on transforming from a first\-person to an allocentric perspective\. This requires the agent not only to comprehend spatial relationships between objects, but also to map these relationships across different viewpoints\. Given the complexity involved, all MLLM\-driven agents perform poorly on this capability, with scores consistently lower than those achieved on SC2 tasks\. This highlights the current limitations of MLLMs in handling dynamic reference frame changes and constructing mental maps\.

![Refer to caption](https://arxiv.org/html/2608.07533v1/x14.png)Figure 9\.Ranking of average scores for embodied spatial cognition capabilities across various agents\.
![Refer to caption](https://arxiv.org/html/2608.07533v1/x15.png)Figure 10\.Quantitative comparison of MetaSpace with existing approaches in SC error detection\.

Human Baseline Performance\.To further contextualize the overall performance of embodied agents in spatial cognition, we invite five human participants \(none of whom have any known spatial cognition impairments\) to complete the same set of test cases, serving as a baseline for comparison\.555Considering the large scale of the test cases, which would create a significant workload for human respondents, we decided to conduct the evaluation on a subset comprising 10% of the original dataset\.As shown in Fig\.[8](https://arxiv.org/html/2608.07533#S6.F8)and Fig\.[8](https://arxiv.org/html/2608.07533#S6.F8), humans consistently outperform all embodied agents across all spatial cognition tasks, underscoring the significant gap between current MLLM\-driven embodied agents and human\-level spatial cognition\.

Ranking of Average Scores\.In Fig\.[10](https://arxiv.org/html/2608.07533#S6.F10), we present the ranking of average scores for embodied spatial cognition capabilities across different embodied agents\. The agent powered by the proprietary model GPT\-5 achieves the highest average score of 0\.52, indicating relatively robust spatial cognition abilities\. Notably, open\-source models \(e\.g\., DeepSeek\-VL2\-small and InternVL3\.5\-8B\) also demonstrate competitive performance, with an average score of 0\.47, surpassing the proprietary model GPT\-4o\. While there are variations in average scores among MLLM\-driven agents, these discrepancies are relatively minor, with all models scoring between 0\.44 and 0\.52\. In contrast, the human baseline achieves an average score of 0\.96, highlighting a substantial gap between embodied agents and human performance\. We further clarify that similar rankings in Fig\.[10](https://arxiv.org/html/2608.07533#S6.F10)result from averaging scores across all SCs\. This reflects the generally limited cognitive capabilities of current agents and highlights the substantial human\-agent performance gap\. However, granular diagnostics for each agent are still evident in Fig\.[8](https://arxiv.org/html/2608.07533#S6.F8), such as identifying GPT\-5’s specific weakness in SC3\-b\. By leveraging MT with fine\-grained MRs to comprehensively assess four SCs, MetaSpace mitigates the “success by coincidence” issue found in existing benchmarks and provides granular diagnostics that binary success metrics in other benchmarks cannot reveal\.

Table 4\.Normalized error statistics\.Err\. Rate: Error Rate \(%, per test case\);Avg\. Err\. Tra\.: Average Errors per Trajectory;Avg\. Err\. Obj\.: Average Errors per Object\.ANSWER to RQ1Our evaluation using MetaSpace reveals that benchmark MLLM\-driven embodied agents exhibit significant limitations in embodied spatial cognition, particularly in tasks involving direction estimation and spatial directional reasoning\. While these agents perform relatively well on magnitude estimation tasks, their overall scores remain substantially lower than human performance\.

### 6\.3\.RQ2: Comparison with Existing Works

#### 6\.3\.1\.Qualitative Analysis

We qualitatively compare MetaSpace with the SOTA embodied agents benchmarks and embodied spatial cognition evaluation approaches to illustrate the advantages of MetaSpace\. As shown in Tab\.[5](https://arxiv.org/html/2608.07533#S6.T5), we compare MetaSpace with EmbodiedBench\(Yang et al\.,[2025](https://arxiv.org/html/2608.07533#bib.bib70)\), 3DSRBench\(Ma et al\.,[2025](https://arxiv.org/html/2608.07533#bib.bib34)\), ECBench\(Dang et al\.,[2025](https://arxiv.org/html/2608.07533#bib.bib16)\), EmbSpatial\-Bench\(Du et al\.,[2024](https://arxiv.org/html/2608.07533#bib.bib18)\), and SPACE\(Ramakrishnan et al\.,[2025](https://arxiv.org/html/2608.07533#bib.bib48)\)from four core dimensions\.

Test Case Generation\.Existing benchmarks primarily rely on manually annotated VQA or MCQ datasets, which are labor\-intensive and costly to create, limiting their scalability\. Besides, the quality of manually created test cases can vary significantly due to human subjectivity, leading to potential biases and inconsistencies\. In contrast, MetaSpace employs an automated test case generation approach based on MRs, enabling the creation of a vast number of test cases without human intervention\. This automation not only enhances scalability but also ensures reproducibility and consistency in test case generation\.

Embodiment\.Some existing benchmarks \(e\.g\., 3DSRBench, EmbSpatial\-Bench\) focus on static VQA tasks without considering the embodied nature of agents, which limits their ability to evaluate spatial cognition in embodied settings\. For example, SPACE includes some classic human cognitive tests \(e\.g\., Minnesota Paper Form Board test\), which do not involve any interaction with the environment\. In contrast, MetaSpace is specifically designed for testing in real embodied settings, by utilizing the real embodied execution trajectories of agents to generate test cases\. Our approach allows for a more realistic embodiment evaluation of spatial cognition capabilities\.

Test Oracle\.Existing benchmarks typically use high\-level task success, VQA/MCQ correctness \(by comparing with human\-annotated answers\), or human evaluation as the test oracle\. However, high\-level task success can miss errors that do not directly cause task failure \(Fig\.[1](https://arxiv.org/html/2608.07533#S1.F1)\), while manual annotation and human evaluation are labor\-intensive and may introduce noise\. In contrast, MetaSpace uses violations of MRs as the test oracle\. These MRs are grounded in established logic rules and physical laws, providing an objective, scalable, and reliable validation process\.

Spatial Cognition Categories\.Existing benchmarks typically emphasize SC2 \(spatial reasoning\), with limited coverage of other essential SC capabilities \(e\.g\., SC1, SC3, and SC4\)\. This narrow focus restricts the ability to holistically evaluate embodied agents\. In contrast, MetaSpace provides comprehensive coverage of all eight key SC capabilities, enabling a more complete and robust assessment of embodied SC\. We will further discuss the necessity of these SCs in §[6\.3\.2](https://arxiv.org/html/2608.07533#S6.SS3.SSS2)\.

Table 5\.Qualitative comparison of MetaSpace with existing SOTA evaluation approaches\.Test Case GenerationEmbodimentTest OracleSC CategoriesMetaSpaceAuto\-gen based on MRs✓MR violationSC1, SC2, SC3, SC4EmbodiedBench\-✓Task successSC2, Perception3DSRBenchMan\. annotated VQAVQA correctnessSC2,Height estimationEmbSpatial\-BenchMCQ correctnessSC2\-a, SC3\-bECBenchMan\. annotated VQA✓Human eval\.HallucinationSPACE∖\\sqrt\{\}\\mkern\-9\.0mu\{\\smallsetminus\}Interactive tasksSPLSC2
#### 6\.3\.2\.Quantitative Analysis

To compare MetaSpace with existing SOTA approaches in testing SC capabilities of embodied agents, we do a small\-scale quantitative analysis\. The challenge is that different methods may target distinct aspects of spatial cognition, making direct quantitative comparison among detection results impractical and unfair\. Therefore, we design a heuristic evaluation protocol to solve this challenge\. Specifically, we apply MetaSpace and other approaches to the trajectories executed by the same embodied agent \(these trajectories are not utilized in our framework to avoid data leakage\), and collect the union of all detected errors\. Next, we invite five independent human reviewers, who have no involvement in the design of our framework and possess extensive experience in embodied intelligence, to annotate whether each detected error is*meaningful*for embodied tasks\. The annotation criteria focus on whether the error could negatively impact embodied task performance\. These negative impacts include: \(1\) errors that directly cause task failure, \(2\) errors that reduce task efficiency, \(3\) errors that may lead to safety risks in real\-world deployment, and \(4\) errors that did not cause failure in the current task due to coincidence but represent potential risks in similar cases\. This protocol allows us to evaluate the proportion of truly*meaningful*spatial cognition errors detected by each method fairly\. Finally, we calculate the ratio of*meaningful*errors detected by each method and show results in[Fig\.10](https://arxiv.org/html/2608.07533#S6.F10)\.

From the results, the union of*meaningful*errors detected by all methods has a total of 1080 errors\. Among them, MetaSpace detects 988*meaningful*errors, accounting for 91\.5% of the union\. This corresponds to a false negative rate of approximately 8\.5%, substantially lower than other SOTA approaches\. The remaining 8\.5% stems from scenarios where the agent’s spatial reasoning isconsistent with MRs yet factually incorrect\(e\.g\., claiming A is left of B and B is right of A while ground truth is A is right of B\)\. Such errors are only detectable by human\-annotated strong oracles\. In contrast, 3DSRBench detects 310*meaningful*errors \(29%\), ECBench detects 407*meaningful*errors \(38%\), EmbSpatial\-Bench detects 300*meaningful*errors \(28%\), and SPACE detects 288*meaningful*errors \(27%\)\. It is notable that EmbodiedBench only detects 35 spatial cognition errors, of which all are*meaningful*, but the total number is very small\. This is because EmbodiedBench primarily focuses on high\-level task success as the test oracle, which overlooks many spatial cognition errors that do not directly lead to task failure\. Crucially, since all approaches were evaluated on identical trajectories, the resulting error concentration in MetaSpace is not due to experimental bias\. Instead, it reflects MetaSpace’s capability to test a broader spectrum of SCs than other SOTAs\. These results demonstrate that MetaSpace significantly outperforms existing approaches in detecting spatial cognition errors that are truly*meaningful*for embodied task performance\. We provide a more detailed comparison with related testing efforts in[Section8](https://arxiv.org/html/2608.07533#S8)\.

ANSWER to RQ2Compared to existing SC evaluation approaches, MetaSpace offers significant advantages in test case generation, embodiment, and test oracle design\. Quantitatively, MetaSpace detects a substantially higher proportion of spatial cognition errors that impact embodied task performance\.

### 6\.4\.RQ3: Internal Evaluation

To investigate the effectiveness of individual MRs in detecting embodied SC errors and to assess the reliability of MetaSpace, we conduct an internal evaluation\. Firstly, to understand the effectiveness of different MRs, we conduct an ablation study in[Section6\.4\.1](https://arxiv.org/html/2608.07533#S6.SS4.SSS1)\. After that, in[Section6\.4\.2](https://arxiv.org/html/2608.07533#S6.SS4.SSS2), we perform a false positive analysis to assess the accuracy of MetaSpace in detecting embodied SC errors\.

![Refer to caption](https://arxiv.org/html/2608.07533v1/x16.png)Figure 11\.MRs that trigger the most SC errors on diverse agents across SCs\. The number in each cell represents the triggered SC error ratio of the corresponding MR type, calculated by\|Errors​\(M​Ru,S​Cv\)\|/\|Total Errors​\(S​Cv\)\|\|\\text\{Errors\}\(MR\_\{u\},SC\_\{v\}\)\|\\,/\\,\|\\text\{Total Errors\}\(SC\_\{v\}\)\|\.
![Refer to caption](https://arxiv.org/html/2608.07533v1/x17.png)Figure 12\.An example of a cognitive map, where the agent predicts object/landmark positions within an 11×\\times11grid11\\text\{\\,\}\\mathrm\{g\}\\mathrm\{r\}\\mathrm\{i\}\\mathrm\{d\}, with the agent located at the center of the map\.

#### 6\.4\.1\.Ablation Study

To assess the effectiveness of different MRs in detecting embodied spatial cognition errors, we conduct an ablation study\. For better visualization and understanding, we present the distribution of embodied spatial cognition errors discovered with various MRs in Fig\.[12](https://arxiv.org/html/2608.07533#S6.F12)\. This figure illustrates which MR is able to identify more spatial cognition errors for different SCs and agents\. The blue section represents the error detection ratio ofMR1,MR2andMR3in directional SCs \(i\.e\., SC1\-a, SC2\-a, SC4\-a\), while the green section shows the number of errors detected byMR2,MR3, andMR4in magnitude SCs \(i\.e\., SC1\-b, SC2\-b, SC4\-b\)\.MR1identifies relatively more spatial cognition errors in directional tasks, whereas MR3 detects a relatively higher number of errors in magnitude SCs\. Overall, all types of MR identify a significant number of spatial cognition errors across various embodied spatial cognition types, so we can conclude that all MRs contribute meaningfully to the overall evaluation of embodied spatial cognition\.

#### 6\.4\.2\.False Positive Analysis

To evaluate the accuracy of MetaSpace in detecting embodied SC errors, we conducted a manual validation of a randomly selected subset of 1,000 detected errors from the total set of violationsV~\\widetilde\{V\}\. The results indicate that MetaSpace produced nine false positives, achieving a precision rate of 99\.1%\. Adjusting for this 0\.9% FP rate shifts overall scores negligibly \(approx\.0\.0050\.005\), leaving our core conclusions unaltered\. Similarly, other SOTAs mentioned in[Section6\.3\.2](https://arxiv.org/html/2608.07533#S6.SS3.SSS2)also suffer from FPs due to annotator inconsistency and subjective bias\.

Notably, all false positive instances were concentrated in the evaluations of MR5 and MR6\. In these two MRs, we employed fixed thresholds \(e\.g\.,ϵ=0\.1\\epsilon=0\.1for MR5 andδ=0\.05\\delta=0\.05for MR6\) to quantify violations, which were derived from human baselines\. However, it is important to recognize that in extreme test cases, human participants often struggle to make precise judgments\. This can lead to larger estimation discrepancies and instances where thresholds exceedϵ=0\.1\\epsilon=0\.1\. While this limitation is acknowledged, it is also a necessary trade\-off to achieve the benefits of quantitative, automated verification\. Overall, the results underscore MetaSpace’s high precision in advancing error detection in embodied spatial cognition\. We discuss the generalizability of this result in[Section3\.4](https://arxiv.org/html/2608.07533#S3.SS4)and further discuss the accuracy in[Section7](https://arxiv.org/html/2608.07533#S7)\.

ANSWER to RQ3Our ablation study reveals that all MRs contribute meaningfully to identifying spatial cognition errors across various SCs, demonstrating their collective effectiveness in evaluating embodied spatial cognition\. Moreover, our false positive analysis on the 1,000 randomly selected errors confirms the high accuracy of MetaSpace, achieving a precision of 99\.1% in error detection\.

### 6\.5\.RQ4: Mitigation

In light of the significant number of embodied spatial cognition errors observed, we conduct case studies to investigate the underlying causes of these errors and explore potential mitigation strategies\. However, given the extensive workload required to examine all spatial cognition types and the fact that this falls somewhatoutside the primary focusof our research \(i\.e\., testing embodied spatial cognition\), we decide to concentrate our efforts on the most critical issue: SC4\-a: directional spatial reasoning under egocentric\-allocentric transformation\. This focus enables us to provide a thorough analysis and develop targeted solutions without being overwhelmed by the breadth of the problem\. Meanwhile, we acknowledge that addressing the broader spectrum of embodied spatial cognition errors will require additional research, which we plan to pursue in future work\. We conduct an error analysis including a case study and propose mitigation strategies as follows\. All experiments are conducted on GPT\-4o\.

#### 6\.5\.1\.Error Analysis

In SC4\-a, we observe that all six benchmark MLLM\-driven embodied agents perform poorly \(Fig\.[8](https://arxiv.org/html/2608.07533#S6.F8)\), with scores below 0\.35\. To investigate the root causes of these errors, we randomly selected 100 detected errors from the total violations for manual analysis\. Our investigation reveals that most of these errors occur during the agents’ reference frame transformations, which involve translating spatial relationships obtained from one reference frame to another\.

In the case illustrated in Fig\.[13](https://arxiv.org/html/2608.07533#S6.F13), the agent is required to determine the direction between two objects \(i\.e\., the kettle and the window\) based on an allocentric description\. Notably, the agent demonstrates an impressive ability to perform step\-by\-step reasoning, outlining processes such as “original orientation”, “from the door”, and “when you turn to” for spatial tasks\. This showcases its capability in modeling spatial reasoning\. Additionally, the agent recognizes the nuances of perspective transformation and differentiates between spatial relationships, indicating its ability to understand egocentric and allocentric perspectives\. Meanwhile, it attempts to transform the reference frame system to infer the answer, rather than resorting to random guessing\.

Initially, the agent accurately describes directions within a single frame of reference, both before and after reference frame transformation\. However, it encounters errors when translating spatial information from one perspective to another, leading to the collapse of the entire reasoning chain\. This indicates that while the agent can perceive spatial relationships within a single reference frame, it struggles to translate these relationships across different perspectives\.

![Refer to caption](https://arxiv.org/html/2608.07533v1/x18.png)Figure 13\.Case study of embodied spatial cognition errors in SC4\-a, usingorangeandgreenbackgrounds to highlight incorrect and correct reasoning steps, respectively\. Note: Structured constraints apply solely to the final action outputs; the intermediate reasoning process via chain of thought \(CoT\) remains natural and unaffected\. This design aligns with the widely adopted “CoT reasoning \+ structured action” paradigm in embodied systems\(Yang et al\.,[2025](https://arxiv.org/html/2608.07533#bib.bib70); Savva et al\.,[2019](https://arxiv.org/html/2608.07533#bib.bib54)\)\.Table 6\.Performance comparison of different prompting techniques in SC4\-a \(measured byscore\)\.
#### 6\.5\.2\.Mitigation Strategies

After identifying the primary root causes of errors in SC4\-a, we propose potential mitigation strategies to address these issues\. The decision\-making of embodied agents is influenced by both the model itself \(e\.g\., MLLM\) and the instruction prompts\. Therefore, similar to performance enhancements in LLMs\(Wu et al\.,[2025](https://arxiv.org/html/2608.07533#bib.bib64); Han et al\.,[2024](https://arxiv.org/html/2608.07533#bib.bib23); Sahoo et al\.,[2024](https://arxiv.org/html/2608.07533#bib.bib53); Chen et al\.,[2025](https://arxiv.org/html/2608.07533#bib.bib9); Liu et al\.,[2023a](https://arxiv.org/html/2608.07533#bib.bib33)\), there is potential to improve the performance of embodied agents through various approaches, including but not limited to model architecture redesign, task\-specific fine\-tuning, and prompt engineering\. In this work, we focus specifically on prompt engineering to mitigate the identified issues in SC4\-a\. We employ two prompting techniques: \(1\) Chain\-of\-Thought \(CoT\) Prompting, which is frequently used in NLP tasks to encourages the model to generate intermediate reasoning steps, thereby enhancing its ability to perform complex reasoning tasks; and \(2\) Cognitive Map Prompting, which asks the MLLM to explicitly construct a mental map of the environment and then use this map to perform spatial tasks\. Fig\.[12](https://arxiv.org/html/2608.07533#S6.F12)illustrates an example of cognitive map\. The results are shown in Tab\.[6](https://arxiv.org/html/2608.07533#S6.T6), where CoT prompting yields a limited improvement, increasing the score from 0\.11 to 0\.15\. In contrast, Cognitive Map Prompting significantly enhances performance, boosting the score to 0\.45\. This substantial improvement suggests that building a mental spatial model or cognitive map serves as a promising solution to tackle SC problems in embodied settings\.

ANSWER to RQ4Our error analysis identifies reference frame transformation as the primary cause of SC4\-a failures\. To mitigate this, we explore prompt engineering and find that cognitive map prompting significantly improves performance by constructing an environmental mental map, whereas traditional techniques \(e\.g\., CoT\) yield only marginal gains\.

## 7\.Discussion

### 7\.1\.Threat to Validity

Internal Validity\.Our manual validation demonstrates high precision \(see[Section6\.4\.2](https://arxiv.org/html/2608.07533#S6.SS4.SSS2)\), with generalizability discussed in[Section3\.4](https://arxiv.org/html/2608.07533#S3.SS4)\. However, we acknowledge that false negatives may still occur\. This limitation is common in testing and highlights the need for ongoing refinement\. Open\-sourced MetaSpace allows the community to contribute additional MRs, thereby improving coverage and reducing false negatives over time\. Furthermore, one might question whether reasoning failures matter if the agent’s final action is correct\. We contend that achieving correct actions from flawed reasoning amounts to “success by coincidence”\. Such models pose safety risks and undermine trust\. MetaSpace specifically exposes these latent cognitive defects that outcome\-oriented metrics miss\.

External Validity\.The current implementation of MetaSpace focuses on three embodied scenarios and eight SC capabilities\. While this selection is diverse, it does not encompass all potential embodied tasks\. Additionally, all experiments are conducted in simulated environments, which may introduce the Sim\-to\-Real Gap\. However, MetaSpace mitigates this concern through its robust design\. The logic\-based MRs rely on qualitative relations \(e\.g\., Transitivity\) that are inherently resilient to pixel\-level noise, while the physics\-based MRs utilize relative ratios and tolerance thresholds to filter out systematic errors\. Our experiments in[Section6](https://arxiv.org/html/2608.07533#S6)focus on RGB\-based MLLMs\. Nonetheless, as discussed in[Section3\.4](https://arxiv.org/html/2608.07533#S3.SS4), our MR design constraints target fundamental attributes of spatial cognition \(e\.g\., direction, distance\), rather than specific SCs or agent types, thereby allowing for broader applicability across various embodiments and agents\. With the open\-sourced MetaSpace framework, we believe researchers can extend the current set of MRs to better meet their specific evaluation needs\. Also, our future work will focus on investigating these validity concerns\.

### 7\.2\.Takeaway Messages

Prioritize Embodied Spatial Cognition\.Enhancing the embodied spatial cognition capabilities of MLLM\-driven embodied agents is essential, particularly in areas such as direction estimation, spatial directional reasoning, and egocentric\-allocentric transformations\. This enhancement spans model architecture design, task\-specific fine\-tuning, and developing self\-supervised learning objectives for spatial reasoning to build robust, embodied\-adapted MLLMs\. Since embodied spatial capabilities are the cornerstone of embodied tasks, pursuing advanced high\-level tasks is futile without first addressing these fundamental spatial cognition challenges\.

Spatially\-Aware Engineering for Embodied Systems\.The limited effectiveness of NLP techniques \(e\.g\., CoT\) in spatial tasks underscores the need for domain\-specific engineering approaches\. Our cognitive map prompting results demonstrate that spatially\-aware interventions can achieve relatively better performance, emphasizing embodied agents demand fundamentally different architectural designs and prompt engineering strategies compared with LLMs\. We advocate for the development of embodied\-specific design patterns, training objectives, and prompting strategies that account for the unique challenges of spatial understanding and environmental interaction\.

## 8\.Related Work

Testing and Evaluation of AI Systems\.Software testing techniques have been widely applied to AI systems, such as deep learning systems\(Pei et al\.,[2017](https://arxiv.org/html/2608.07533#bib.bib46); Wang and Su,[2020](https://arxiv.org/html/2608.07533#bib.bib61); Yuan et al\.,[2021](https://arxiv.org/html/2608.07533#bib.bib72)\), autonomous driving systems\(Zhang et al\.,[2018](https://arxiv.org/html/2608.07533#bib.bib74); Tian et al\.,[2018](https://arxiv.org/html/2608.07533#bib.bib56)\), and LLMs\(Zhang et al\.,[2024b](https://arxiv.org/html/2608.07533#bib.bib77); Bouzenia et al\.,[2025](https://arxiv.org/html/2608.07533#bib.bib5); Li et al\.,[2024a](https://arxiv.org/html/2608.07533#bib.bib29)\)\. However, these approaches do not address the unique challenges of embodied agents\. First, existing frameworks primarily evaluate static inputs \(e\.g\., single images or text prompts\), whereas embodied agents operate in dynamic environments\. Embodied tasks require agents to reason about continuous state transitions rather than isolated snapshots\. MetaSpace evaluates the logical coherence of the agent’s spatial understanding across temporal sequences \(e\.g\., during navigation or manipulation\), a dimension largely absent in existing testing tools\. Secondly, traditional MT often relies on semantic invariance \(e\.g\., robustness against pixel noise or synonym substitution\)\. MetaSpace, however, introduces physical and logical covariance\. We define MRs based on logical rules and physical laws, requiring the agent’s reasoning to evolve consistently with its physical interactions, rather than simply remaining invariant to perturbations\. Thirdly, while some Video Question Answering benchmarks\(Zhang et al\.,[2025a](https://arxiv.org/html/2608.07533#bib.bib76),[b](https://arxiv.org/html/2608.07533#bib.bib78)\)address dynamic features, they focus on semantic understanding via passive perception\. In contrast, MetaSpace targets the active perception\-action loop\. We verify whether the agent constructs a consistent internal mental map \(SC4\) and correctly perceives its own ego\-motion \(SC1\) to guide actions\. This shifts the evaluation focus from passive pattern recognition to active spatial cognition\.

Neurosymbolic Approaches\.Recent work in neurosymbolic has focused on integrating symbolic operators with neural perception modules to combine their complementary strengths\(Andreas et al\.,[2016](https://arxiv.org/html/2608.07533#bib.bib2); Yi et al\.,[2018](https://arxiv.org/html/2608.07533#bib.bib71); Verbruggen et al\.,[2021](https://arxiv.org/html/2608.07533#bib.bib59); Li et al\.,[2023](https://arxiv.org/html/2608.07533#bib.bib30)\)\. These approaches have led to novel solutions across several domains, including the synthesis of neurosymbolic programs for fine\-grained image editing\(Barnaby et al\.,[2023](https://arxiv.org/html/2608.07533#bib.bib4)\)or image interpretation\(Mao et al\.,[2019](https://arxiv.org/html/2608.07533#bib.bib36)\), the creation of semantic regular expressions for data extraction\(Chen et al\.,[2023](https://arxiv.org/html/2608.07533#bib.bib10)\), and the generation of executable programs from natural language questions\(Chen et al\.,[2021](https://arxiv.org/html/2608.07533#bib.bib11)\)\. Unlike these works, which focus on building integrated reasoning systems, our approach applies neurosymbolic principles to the testing and validation phase\. Specifically, MetaSpace uses symbolic MRs to create formal and precise specifications for the ambiguous spatial cognition concept of embodied agents\.

Property\-Based Testing\.Property\-based testing \(PBT\)\(Claessen and Hughes,[2000](https://arxiv.org/html/2608.07533#bib.bib15)\)is a mainstream technique that verifies high\-level properties of a system against a multitude of randomly generated inputs\. Prior work has applied PBT to test a wide range of systems\(Goldstein et al\.,[2024](https://arxiv.org/html/2608.07533#bib.bib22)\), from validating compilers for languages like Haskell\(Pałka et al\.,[2011](https://arxiv.org/html/2608.07533#bib.bib44)\)and OCaml\(Midtgaard et al\.,[2017](https://arxiv.org/html/2608.07533#bib.bib38)\)to automated testing of Android applications\(Xiong et al\.,[2024](https://arxiv.org/html/2608.07533#bib.bib68)\)and performing acceptance testing for complex web user interfaces\(O’Connor and Wickström,[2022](https://arxiv.org/html/2608.07533#bib.bib41)\)\. Our work extends this paradigm to the fundamentally different domain of embodied spatial cognition, an area previously unexplored by PBT\. The core contribution lies in crafting novel MRs grounded not in program semantics, but in the principles of logic and physics that govern an agent’s interaction with its environment\. These MRs function ascognitive invariants, formalizing and testing the underlying spatial cognition properties of embodied agents in a continuous, dynamic environment\.

## 9\.Conclusion

We present MetaSpace, a metamorphic testing framework for embodied spatial cognition, revealing significant limitations in current agents\. We show that cognitive map prompting mitigates these errors, advocating for improved spatial cognition to build reliable real\-world agents\.

## Data\-Availability Statement

The MetaSpace framework, including source code, experiment scripts, and reproduction instructions, is available in the artifact\(Xu et al\.,[2026](https://arxiv.org/html/2608.07533#bib.bib69)\)\. Due to privacy and licensing constraints, some third\-party datasets used in our experiments \(e\.g\., EB\-Navigation, EB\-Manipulation, AerialVLN\) cannot be redistributed directly\. However, we provide documentation to assist users in obtaining these datasets from their official sources\. All custom\-generated configuration files are included in the artifact package, designed to facilitate reproducibility and further research\.

Please note that MLLM inference requires access to proprietary model APIs, open\-source model deployments, and suitable hardware \(e\.g\., GPU\) for full\-scale experiments, which are not included in the artifact package\. Again, we provide relevant links to guide users in obtaining access to these resources\. Additional limitations and setup instructions are documented in the artifact repository\.

###### Acknowledgements\.

The HKUST authors were supported in part by a RGC GRF grant under the contract 16214723, an ITF grant under the contract ITS/161/24FP, and a HKUST Bridge The Gap fund BGF\.001\.2025\. We are grateful to the anonymous reviewers for their valuable comments\.

## References

- \(1\)
- Andreas et al\.\(2016\)Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein\. 2016\.Learning to Compose Neural Networks for Question Answering\. In*Proceedings of NAACL\-HLT*\. 1545–1554\.[doi:10\.18653/v1/N16\-1181](https://doi.org/10.18653/v1/N16-1181)
- Anthropic \(2025\)Anthropic\. 2025\.Claude Sonnet 4\.[https://www\.anthropic\.com/claude/sonnet](https://www.anthropic.com/claude/sonnet)
- Barnaby et al\.\(2023\)Celeste Barnaby, Qiaochu Chen, Roopsha Samanta, and Işıl Dillig\. 2023\.Imageeye: Batch image processing using program synthesis\.*Proceedings of the ACM on Programming Languages*7, PLDI \(2023\), 686–711\.[doi:10\.1145/3591248](https://doi.org/10.1145/3591248)
- Bouzenia et al\.\(2025\)Islem Bouzenia, Premkumar Devanbu, and Michael Pradel\. 2025\.RepairAgent: An Autonomous, LLM\-Based Agent for Program Repair\. In*2025 IEEE/ACM 47th International Conference on Software Engineering \(ICSE\)*\. IEEE, 2188–2200\.[doi:10\.1109/ICSE55347\.2025\.00157](https://doi.org/10.1109/ICSE55347.2025.00157)
- Burgess \(2006\)Neil Burgess\. 2006\.Spatial memory: how egocentric and allocentric combine\.*Trends in cognitive sciences*10, 12 \(2006\), 551–557\.[doi:10\.1016/j\.tics\.2006\.10\.005](https://doi.org/10.1016/j.tics.2006.10.005)
- Burgess \(2008\)Neil Burgess\. 2008\.Spatial cognition and the brain\.*Annals of the New York Academy of Sciences*1124, 1 \(2008\), 77–97\.[doi:10\.1196/annals\.1440\.002](https://doi.org/10.1196/annals.1440.002)
- Chatterjee et al\.\(2024\)Agneet Chatterjee, Gabriela Ben Melech Stan, Estelle Aflalo, Sayak Paul, Dhruba Ghosh, Tejas Gokhale, Ludwig Schmidt, Hannaneh Hajishirzi, Vasudev Lal, Chitta Baral, et al\.2024\.Getting it right: Improving spatial consistency in text\-to\-image models\. In*European Conference on Computer Vision*\. Springer, 204–222\.[doi:10\.1007/978\-3\-031\-72670\-5\_12](https://doi.org/10.1007/978-3-031-72670-5_12)
- Chen et al\.\(2025\)Banghao Chen, Zhaofeng Zhang, Nicolas Langrené, and Shengxin Zhu\. 2025\.Unleashing the potential of prompt engineering for large language models\.*Patterns*\(2025\)\.[doi:10\.1016/j\.patter\.2025\.101260](https://doi.org/10.1016/j.patter.2025.101260)
- Chen et al\.\(2023\)Qiaochu Chen, Arko Banerjee, Çağatay Demiralp, Greg Durrett, and Işıl Dillig\. 2023\.Data extraction via semantic regular expression synthesis\.*Proceedings of the ACM on Programming Languages*7, OOPSLA2 \(2023\), 1848–1877\.[doi:10\.1145/3622863](https://doi.org/10.1145/3622863)
- Chen et al\.\(2021\)Qiaochu Chen, Aaron Lamoreaux, Xinyu Wang, Greg Durrett, Osbert Bastani, and Isil Dillig\. 2021\.Web question answering with neurosymbolic program synthesis\. In*Proceedings of the 42nd ACM SIGPLAN International Conference on Programming Language Design and Implementation*\. 328–343\.[doi:10\.1145/3453483\.3454047](https://doi.org/10.1145/3453483.3454047)
- Chen et al\.\(2020\)Tsong Y Chen, Shing C Cheung, and Shiu Ming Yiu\. 2020\.Metamorphic testing: a new approach for generating next test cases\.*arXiv preprint arXiv:2002\.12543*\(2020\)\.
- Cheng et al\.\(2025\)Zhili Cheng, Yuge Tu, Ran Li, Shiqi Dai, Jinyi Hu, Shengding Hu, Jiahao Li, Yang Shi, Tianyu Yu, Weize Chen, et al\.2025\.Embodiedeval: Evaluate multimodal llms as embodied agents\.*arXiv preprint arXiv:2501\.11858*\(2025\)\.
- Choi et al\.\(2024\)Jae\-Woo Choi, Youngwoo Yoon, Hyobin Ong, Jaehong Kim, and Minsu Jang\. 2024\.LoTa\-Bench: Benchmarking Language\-oriented Task Planners for Embodied Agents\. In*International Conference on Learning Representations \(ICLR\) 2024*\. 1–27\.[https://openreview\.net/pdf?id=ADSxCpCu9s](https://openreview.net/pdf?id=ADSxCpCu9s)
- Claessen and Hughes \(2000\)Koen Claessen and John Hughes\. 2000\.QuickCheck: a lightweight tool for random testing of Haskell programs\. In*Proceedings of the fifth ACM SIGPLAN international conference on Functional programming*\. 268–279\.[doi:10\.1145/351240\.351266](https://doi.org/10.1145/351240.351266)
- Dang et al\.\(2025\)Ronghao Dang, Yuqian Yuan, Wenqi Zhang, Yifei Xin, Boqiang Zhang, Long Li, Liuyi Wang, Qinyang Zeng, Xin Li, and Lidong Bing\. 2025\.ECBench: Can Multi\-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition Benchmark\. In*Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\)*\. 24593–24602\.[doi:10\.1109/CVPR52734\.2025\.02290](https://doi.org/10.1109/CVPR52734.2025.02290)
- Dong et al\.\(2025\)Jinhao Dong, Jun Sun, Wenjie Zhang, Jin Song Dong, and Dan Hao\. 2025\.ConTested: Consistency\-Aided Tested Code Generation with LLM\.*Proceedings of the ACM on Software Engineering*2, ISSTA \(2025\), 596–617\.[doi:10\.1145/3728902](https://doi.org/10.1145/3728902)
- Du et al\.\(2024\)Mengfei Du, Binhao Wu, Zejun Li, Xuan\-Jing Huang, and Zhongyu Wei\. 2024\.Embspatial\-bench: Benchmarking spatial understanding for embodied tasks with large vision\-language models\. In*Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 2: Short Papers\)*\. 346–355\.[doi:10\.18653/v1/2024\.acl\-short\.33](https://doi.org/10.18653/v1/2024.acl-short.33)
- Friedman \(1937\)Milton Friedman\. 1937\.The use of ranks to avoid the assumption of normality implicit in the analysis of variance\.*Journal of the american statistical association*32, 200 \(1937\), 675–701\.[doi:10\.1080/01621459\.1937\.10503522](https://doi.org/10.1080/01621459.1937.10503522)
- Gao et al\.\(2024\)Chen Gao, Baining Zhao, Weichen Zhang, Jinzhu Mao, Jun Zhang, Zhiheng Zheng, Fanhang Man, Jianjie Fang, Zile Zhou, Jinqiang Cui, et al\.2024\.Embodiedcity: A benchmark platform for embodied agent in real\-world city environment\.*arXiv preprint arXiv:2410\.09604*\(2024\)\.
- Gibson \(2014\)James J Gibson\. 2014\.*The ecological approach to visual perception: classic edition*\.Psychology press\.[doi:10\.4324/9781315740218](https://doi.org/10.4324/9781315740218)
- Goldstein et al\.\(2024\)Harrison Goldstein, Joseph W Cutler, Daniel Dickstein, Benjamin C Pierce, and Andrew Head\. 2024\.Property\-based testing in practice\. In*Proceedings of the IEEE/ACM 46th International Conference on Software Engineering*\. 1–13\.[doi:10\.1145/3597503\.3639581](https://doi.org/10.1145/3597503.3639581)
- Han et al\.\(2024\)Zeyu Han, Chao Gao, Jinyang Liu, Jeff Zhang, and Sai Qian Zhang\. 2024\.Parameter\-efficient fine\-tuning for large models: A comprehensive survey\.*arXiv preprint arXiv:2403\.14608*\(2024\)\.
- Hoehing et al\.\(2023\)Nils Hoehing, Ellen Rushe, and Anthony Ventresque\. 2023\.What’s left can’t be right–The remaining positional incompetence of contrastive vision\-language models\.*arXiv preprint arXiv:2311\.11477*\(2023\)\.
- Kolve et al\.\(2017\)Eric Kolve, Roozbeh Mottaghi, Winson Han, Eli VanderBilt, Luca Weihs, Alvaro Herrasti, Matt Deitke, Kiana Ehsani, Daniel Gordon, Yuke Zhu, et al\.2017\.Ai2\-thor: An interactive 3d environment for visual ai\.*arXiv preprint arXiv:1712\.05474*\(2017\)\.
- Kong et al\.\(2024\)Xiangrui Kong, Wenxiao Zhang, Jin Hong, and Thomas Braunl\. 2024\.Embodied AI in mobile robots: Coverage path planning with large language models\.*arXiv preprint arXiv:2407\.02220*\(2024\)\.
- Kosslyn \(1987\)Stephen M Kosslyn\. 1987\.Seeing and imagining in the cerebral hemispheres: a computational approach\.*Psychological review*94, 2 \(1987\), 148\.[doi:10\.1037/0033\-295X\.94\.2\.148](https://doi.org/10.1037/0033-295X.94.2.148)
- Li et al\.\(2024b\)Manling Li, Shiyu Zhao, Qineng Wang, Kangrui Wang, Yu Zhou, Sanjana Srivastava, Cem Gokmen, Tony Lee, Li Erran Li, Ruohan Zhang, et al\.2024b\.Embodied agent interface: benchmarking LLMs for embodied decision making\. In*Proceedings of the 38th International Conference on Neural Information Processing Systems*\. 100428–100534\.[doi:10\.52202/079017\-3188](https://doi.org/10.52202/079017-3188)
- Li et al\.\(2024a\)Ningke Li, Yuekang Li, Yi Liu, Ling Shi, Kailong Wang, and Haoyu Wang\. 2024a\.Drowzee: Metamorphic testing for fact\-conflicting hallucination detection in large language models\.*Proceedings of the ACM on Programming Languages*8, OOPSLA2 \(2024\), 1843–1872\.[doi:10\.1145/3689776](https://doi.org/10.1145/3689776)
- Li et al\.\(2023\)Ziyang Li, Jiani Huang, and Mayur Naik\. 2023\.Scallop: A language for neurosymbolic programming\.*Proceedings of the ACM on Programming Languages*7, PLDI \(2023\), 1463–1487\.[doi:10\.1145/3591280](https://doi.org/10.1145/3591280)
- Likert and Quasha \(1941\)Rensis Likert and WH Quasha\. 1941\.*Minnesota Paper Form Board Test*\.Psychological Corporation\.
- Liu et al\.\(2023b\)Shubo Liu, Hongsheng Zhang, Yuankai Qi, Peng Wang, Yanning Zhang, and Qi Wu\. 2023b\.Aerialvln: Vision\-and\-language navigation for uavs\. In*Proceedings of the IEEE/CVF International Conference on Computer Vision*\. 15384–15394\.[doi:10\.1109/ICCV51070\.2023\.01411](https://doi.org/10.1109/ICCV51070.2023.01411)
- Liu et al\.\(2023a\)Xiaoxia Liu, Jingyi Wang, Jun Sun, Xiaohan Yuan, Guoliang Dong, Peng Di, Wenhai Wang, and Dongxia Wang\. 2023a\.Prompting frameworks for large language models: A survey\.*arXiv preprint arXiv:2311\.12785*\(2023\)\.
- Ma et al\.\(2025\)Wufei Ma, Haoyu Chen, Guofeng Zhang, Yu\-Cheng Chou, Jieneng Chen, Celso de Melo, and Alan Yuille\. 2025\.3dsrbench: A comprehensive 3d spatial reasoning benchmark\. In*Proceedings of the IEEE/CVF International Conference on Computer Vision*\. 6924–6934\.[https://openaccess\.thecvf\.com/content/ICCV2025/html/Ma\_3DSRBench\_A\_Comprehensive\_3D\_Spatial\_Reasoning\_Benchmark\_ICCV\_2025\_paper\.html](https://openaccess.thecvf.com/content/ICCV2025/html/Ma_3DSRBench_A_Comprehensive_3D_Spatial_Reasoning_Benchmark_ICCV_2025_paper.html)
- Mansur et al\.\(2021\)Muhammad Numair Mansur, Maria Christakis, and Valentin Wüstholz\. 2021\.Metamorphic testing of Datalog engines\. In*Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering*\. 639–650\.[doi:10\.1145/3468264\.3468573](https://doi.org/10.1145/3468264.3468573)
- Mao et al\.\(2019\)Jiayuan Mao, Chuang Gan, Pushmeet Kohli, Joshua B Tenenbaum, and Jiajun Wu\. 2019\.The Neuro\-Symbolic Concept Learner: Interpreting Scenes, Words, and Sentences From Natural Supervision\. In*International Conference on Learning Representations \(ICLR\)*\.[https://openreview\.net/forum?id=rJgMlhRctm](https://openreview.net/forum?id=rJgMlhRctm)
- Marr \(2010\)David Marr\. 2010\.*Vision: A computational investigation into the human representation and processing of visual information*\.MIT press\.[doi:10\.7551/mitpress/9780262514620\.001\.0001](https://doi.org/10.7551/mitpress/9780262514620.001.0001)
- Midtgaard et al\.\(2017\)Jan Midtgaard, Mathias Nygaard Justesen, Patrick Kasting, Flemming Nielson, and Hanne Riis Nielson\. 2017\.Effect\-driven QuickChecking of compilers\.*Proceedings of the ACM on Programming Languages*1, ICFP \(2017\), 1–23\.[doi:10\.1145/3110259](https://doi.org/10.1145/3110259)
- Montello \(2001\)D\.R\. Montello\. 2001\.Spatial Cognition\.In*International Encyclopedia of the Social & Behavioral Sciences*, Neil J\. Smelser and Paul B\. Baltes \(Eds\.\)\. Pergamon, Oxford, 14771–14775\.[doi:10\.1016/B0\-08\-043076\-7/02492\-X](https://doi.org/10.1016/B0-08-043076-7/02492-X)
- Naveed et al\.\(2025\)Humza Naveed, Asad Ullah Khan, Shi Qiu, Muhammad Saqib, Saeed Anwar, Muhammad Usman, Naveed Akhtar, Nick Barnes, and Ajmal Mian\. 2025\.A comprehensive overview of large language models\.*ACM Transactions on Intelligent Systems and Technology*16, 5 \(2025\), 1–72\.[doi:10\.1145/3744746](https://doi.org/10.1145/3744746)
- O’Connor and Wickström \(2022\)Liam O’Connor and Oskar Wickström\. 2022\.Quickstrom: property\-based acceptance testing with LTL specifications\. In*Proceedings of the 43rd ACM SIGPLAN International Conference on Programming Language Design and Implementation*\. 1025–1038\.[doi:10\.1145/3519939\.3523728](https://doi.org/10.1145/3519939.3523728)
- OpenAI \(2024\)OpenAI\. 2024\.Hello GPT\-4o\.[https://openai\.com/index/hello\-gpt\-4o/](https://openai.com/index/hello-gpt-4o/)
- OpenAI \(2025\)OpenAI\. 2025\.Introducing GPT\-5\.[https://openai\.com/index/introducing\-gpt\-5/](https://openai.com/index/introducing-gpt-5/)
- Pałka et al\.\(2011\)Michał H Pałka, Koen Claessen, Alejandro Russo, and John Hughes\. 2011\.Testing an optimising compiler by generating random lambda terms\. In*Proceedings of the 6th International Workshop on Automation of Software Test*\. 91–97\.[doi:10\.1145/1982595\.1982615](https://doi.org/10.1145/1982595.1982615)
- Paltenghi and Pradel \(2023\)Matteo Paltenghi and Michael Pradel\. 2023\.MorphQ: Metamorphic testing of the Qiskit quantum computing platform\. In*2023 IEEE/ACM 45th International Conference on Software Engineering \(ICSE\)*\. IEEE, 2413–2424\.[doi:10\.1109/ICSE48619\.2023\.00202](https://doi.org/10.1109/ICSE48619.2023.00202)
- Pei et al\.\(2017\)Kexin Pei, Yinzhi Cao, Junfeng Yang, and Suman Jana\. 2017\.Deepxplore: Automated whitebox testing of deep learning systems\. In*proceedings of the 26th Symposium on Operating Systems Principles*\. 1–18\.[doi:10\.1145/3132747\.3132785](https://doi.org/10.1145/3132747.3132785)
- Raiaan et al\.\(2024\)Mohaimenul Azam Khan Raiaan, Md Saddam Hossain Mukta, Kaniz Fatema, Nur Mohammad Fahad, Sadman Sakib, Most Marufatul Jannat Mim, Jubaer Ahmad, Mohammed Eunus Ali, and Sami Azam\. 2024\.A review on large language models: Architectures, applications, taxonomies, open issues and challenges\.*IEEE access*12 \(2024\), 26839–26874\.[doi:10\.1109/ACCESS\.2024\.3365742](https://doi.org/10.1109/ACCESS.2024.3365742)
- Ramakrishnan et al\.\(2025\)Santhosh Kumar Ramakrishnan, Erik Wijmans, Philipp Kraehenbuehl, and Vladlen Koltun\. 2025\.Does Spatial Cognition Emerge in Frontier Models?\. In*The Thirteenth International Conference on Learning Representations \(ICLR\)*\.[https://openreview\.net/forum?id=WK6K1FMEQ1](https://openreview.net/forum?id=WK6K1FMEQ1)
- Redmon et al\.\(2016\)Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi\. 2016\.You only look once: Unified, real\-time object detection\. In*Proceedings of the IEEE conference on computer vision and pattern recognition*\. 779–788\.[doi:10\.1109/CVPR\.2016\.91](https://doi.org/10.1109/CVPR.2016.91)
- Rohmer et al\.\(2013\)Eric Rohmer, Surya PN Singh, and Marc Freese\. 2013\.V\-REP: A versatile and scalable robot simulation framework\. In*2013 IEEE/RSJ international conference on intelligent robots and systems*\. IEEE, 1321–1326\.[doi:10\.1109/IROS\.2013\.6696520](https://doi.org/10.1109/IROS.2013.6696520)
- Rostam et al\.\(2024\)Zhyar Rzgar K Rostam, Sándor Szénási, and Gábor Kertész\. 2024\.Achieving peak performance for large language models: A systematic review\.*IEEE access*\(2024\)\.[doi:10\.1109/ACCESS\.2024\.3424945](https://doi.org/10.1109/ACCESS.2024.3424945)
- Roy et al\.\(2021\)Nicholas Roy, Ingmar Posner, Tim Barfoot, Philippe Beaudoin, Yoshua Bengio, Jeannette Bohg, Oliver Brock, Isabelle Depatie, Dieter Fox, Dan Koditschek, et al\.2021\.From machine learning to robotics: Challenges and opportunities for embodied intelligence\.*arXiv preprint arXiv:2110\.15245*\(2021\)\.
- Sahoo et al\.\(2024\)Pranab Sahoo, Ayush Kumar Singh, Sriparna Saha, Vinija Jain, Samrat Mondal, and Aman Chadha\. 2024\.A systematic survey of prompt engineering in large language models: Techniques and applications\.*arXiv preprint arXiv:2402\.07927*\(2024\)\.
- Savva et al\.\(2019\)Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, et al\.2019\.Habitat: A platform for embodied ai research\. In*Proceedings of the IEEE/CVF international conference on computer vision*\. 9339–9347\.[doi:10\.1109/ICCV\.2019\.00943](https://doi.org/10.1109/ICCV.2019.00943)
- Tian et al\.\(2025\)Xiaoyu Tian, Junru Gu, Bailin Li, Yicheng Liu, Yang Wang, Zhiyong Zhao, Kun Zhan, Peng Jia, XianPeng Lang, and Hang Zhao\. 2025\.DriveVLM: The Convergence of Autonomous Driving and Large Vision\-Language Models\. In*Conference on Robot Learning*\. PMLR, 4698–4726\.[https://proceedings\.mlr\.press/v270/tian25c\.html](https://proceedings.mlr.press/v270/tian25c.html)
- Tian et al\.\(2018\)Yuchi Tian, Kexin Pei, Suman Jana, and Baishakhi Ray\. 2018\.Deeptest: Automated testing of deep\-neural\-network\-driven autonomous cars\. In*Proceedings of the 40th international conference on software engineering*\. 303–314\.[doi:10\.1145/3180155\.3180220](https://doi.org/10.1145/3180155.3180220)
- Tolksdorf et al\.\(2019\)Sandro Tolksdorf, Daniel Lehmann, and Michael Pradel\. 2019\.Interactive metamorphic testing of debuggers\. In*Proceedings of the 28th ACM SIGSOFT International Symposium on Software Testing and Analysis*\. 273–283\.[doi:10\.1145/3293882\.3330567](https://doi.org/10.1145/3293882.3330567)
- Ultralytics \(2024\)Ultralytics\. 2024\.Ultralytics YOLO11 Documentation\.[https://docs\.ultralytics\.com/models/yolo11/](https://docs.ultralytics.com/models/yolo11/)
- Verbruggen et al\.\(2021\)Gust Verbruggen, Vu Le, and Sumit Gulwani\. 2021\.Semantic programming by example with pre\-trained models\.*Proceedings of the ACM on Programming Languages*5, OOPSLA \(2021\), 1–25\.[doi:10\.1145/3485477](https://doi.org/10.1145/3485477)
- Wang et al\.\(2024\)Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al\.2024\.Qwen2\-vl: Enhancing vision\-language model’s perception of the world at any resolution\.*arXiv preprint arXiv:2409\.12191*\(2024\)\.
- Wang and Su \(2020\)Shuai Wang and Zhendong Su\. 2020\.Metamorphic object insertion for testing object detection systems\. In*Proceedings of the 35th IEEE/ACM International Conference on Automated Software Engineering*\. 1053–1065\.[doi:10\.1145/3324884\.3416584](https://doi.org/10.1145/3324884.3416584)
- Wang et al\.\(2025\)Weiyun Wang, Zhangwei Gao, Lixin Gu, Hengjun Pu, Long Cui, Xingguang Wei, Zhaoyang Liu, Linglin Jing, Shenglong Ye, Jie Shao, et al\.2025\.Internvl3\. 5: Advancing open\-source multimodal models in versatility, reasoning, and efficiency\.*arXiv preprint arXiv:2508\.18265*\(2025\)\.
- Wielemaker et al\.\(2012\)Jan Wielemaker, Tom Schrijvers, Markus Triska, and Torbjörn Lager\. 2012\.SWI\-Prolog\.*Theory and Practice of Logic Programming*12, 1\-2 \(2012\), 67–96\.[doi:10\.1017/S1471068411000494](https://doi.org/10.1017/S1471068411000494)
- Wu et al\.\(2025\)Xiao\-Kun Wu, Min Chen, Wanyi Li, Rui Wang, Limeng Lu, Jia Liu, Kai Hwang, Yixue Hao, Yanru Pan, Qingguo Meng, et al\.2025\.Llm fine\-tuning: Concepts, opportunities, and challenges\.*Big Data and Cognitive Computing*9, 4 \(2025\), 87\.[doi:10\.3390/bdcc9040087](https://doi.org/10.3390/bdcc9040087)
- Wu et al\.\(2024\)Zhiyu Wu, Xiaokang Chen, Zizheng Pan, Xingchao Liu, Wen Liu, Damai Dai, Huazuo Gao, Yiyang Ma, Chengyue Wu, Bingxuan Wang, et al\.2024\.Deepseek\-vl2: Mixture\-of\-experts vision\-language models for advanced multimodal understanding\.*arXiv preprint arXiv:2412\.10302*\(2024\)\.
- Xiao et al\.\(2025\)Dongwei Xiao, Zhibo Liu, Yiteng Peng, and Shuai Wang\. 2025\.MTZK: Testing and Exploring Bugs in Zero\-Knowledge \(ZK\) Compilers\.\. In*NDSS*\.[doi:10\.14722/ndss\.2025\.230530](https://doi.org/10.14722/ndss.2025.230530)
- Xiao et al\.\(2022\)Dongwei Xiao, Zhibo Liu, Yuanyuan Yuan, Qi Pang, and Shuai Wang\. 2022\.Metamorphic testing of deep learning compilers\.*Proceedings of the ACM on Measurement and Analysis of Computing Systems*6, 1 \(2022\), 1–28\.[doi:10\.1145/3508035](https://doi.org/10.1145/3508035)
- Xiong et al\.\(2024\)Yiheng Xiong, Ting Su, Jue Wang, Jingling Sun, Geguang Pu, and Zhendong Su\. 2024\.General and Practical Property\-based Testing for Android Apps\. In*Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering*\. 53–64\.[doi:10\.1145/3691620\.3694986](https://doi.org/10.1145/3691620.3694986)
- Xu et al\.\(2026\)Gengyang Xu, Dongwei Xiao, Yiteng Peng, and Shuai Wang\. 2026\.Reproduction Package for Article ‘Metaspace: Metamorphic Testing for Spatial Cognition in Embodied Agents’\.Zenodo\.[doi:10\.5281/zenodo\.19394476](https://doi.org/10.5281/zenodo.19394476)
- Yang et al\.\(2025\)Rui Yang, Hanyang Chen, Junyu Zhang, Mark Zhao, Cheng Qian, Kangrui Wang, Qineng Wang, Teja Venkat Koripella, Marziyeh Movahedi, Manling Li, et al\.2025\.EmbodiedBench: Comprehensive Benchmarking Multi\-modal Large Language Models for Vision\-Driven Embodied Agents\. In*International Conference on Machine Learning*\. PMLR, 70576–70631\.[https://proceedings\.mlr\.press/v267/yang25f\.html](https://proceedings.mlr.press/v267/yang25f.html)
- Yi et al\.\(2018\)Kexin Yi, Jiajun Wu, Chuang Gan, Antonio Torralba, Pushmeet Kohli, and Josh Tenenbaum\. 2018\.Neural\-symbolic vqa: Disentangling reasoning from vision and language understanding\.*Advances in neural information processing systems*31 \(2018\)\.[https://dl\.acm\.org/doi/10\.5555/3326943\.3327039](https://dl.acm.org/doi/10.5555/3326943.3327039)
- Yuan et al\.\(2021\)Yuanyuan Yuan, Shuai Wang, Mingyue Jiang, and Tsong Yueh Chen\. 2021\.Perception matters: Detecting perception failures of vqa models using metamorphic testing\. In*Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition*\. 16908–16917\.[doi:10\.1109/CVPR46437\.2021\.01663](https://doi.org/10.1109/CVPR46437.2021.01663)
- Zhang et al\.\(2025c\)Huanyu Zhang, Chengzu Li, Wenshan Wu, Shaoguang Mao, Yifan Zhang, Haochen Tian, Ivan Vulić, Zhang Zhang, Liang Wang, Tieniu Tan, et al\.2025c\.Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes\.*arXiv preprint arXiv:2504\.15037*\(2025\)\.
- Zhang et al\.\(2018\)Mengshi Zhang, Yuqun Zhang, Lingming Zhang, Cong Liu, and Sarfraz Khurshid\. 2018\.Deeproad: Gan\-based metamorphic testing and input validation framework for autonomous driving systems\. In*Proceedings of the 33rd ACM/IEEE international conference on automated software engineering*\. 132–142\.[doi:10\.1145/3238147\.3238187](https://doi.org/10.1145/3238147.3238187)
- Zhang et al\.\(2024a\)Xiao\-Yi Zhang, Yang Liu, Paolo Arcaini, Mingyue Jiang, and Zheng Zheng\. 2024a\.Met\-mapf: A metamorphic testing approach for multi\-agent path finding algorithms\.*ACM Transactions on Software Engineering and Methodology*33, 8 \(2024\), 1–37\.[doi:10\.1145/3669663](https://doi.org/10.1145/3669663)
- Zhang et al\.\(2025a\)Yuanhan Zhang, Yunice Chew, Yuhao Dong, Aria Leo, Bo Hu, and Ziwei Liu\. 2025a\.Towards video thinking test: A holistic benchmark for advanced video reasoning and understanding\. In*Proceedings of the IEEE/CVF International Conference on Computer Vision*\. 20626–20636\.[https://openaccess\.thecvf\.com/content/ICCV2025/papers/Zhang\_Towards\_Video\_Thinking\_Test\_A\_Holistic\_Benchmark\_for\_Advanced\_Video\_ICCV\_2025\_paper\.pdf](https://openaccess.thecvf.com/content/ICCV2025/papers/Zhang_Towards_Video_Thinking_Test_A_Holistic_Benchmark_for_Advanced_Video_ICCV_2025_paper.pdf)
- Zhang et al\.\(2024b\)Yuntong Zhang, Haifeng Ruan, Zhiyu Fan, and Abhik Roychoudhury\. 2024b\.Autocoderover: Autonomous program improvement\. In*Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis*\. 1592–1604\.[doi:10\.1145/3650212\.3680384](https://doi.org/10.1145/3650212.3680384)
- Zhang et al\.\(2025b\)Zicheng Zhang, Ziheng Jia, Haoning Wu, Chunyi Li, Zijian Chen, Yingjie Zhou, Wei Sun, Xiaohong Liu, Xiongkuo Min, Weisi Lin, et al\.2025b\.Q\-Bench\-Video: Benchmark the Video Quality Understanding of LMMs\. In*Proceedings of the Computer Vision and Pattern Recognition Conference*\. 3229–3239\.[doi:10\.1109/CVPR52734\.2025\.00307](https://doi.org/10.1109/CVPR52734.2025.00307)
- Zhao et al\.\(2025\)Baining Zhao, Jianjie Fang, Zichao Dai, Ziyou Wang, Jirong Zha, Weichen Zhang, Chen Gao, Yue Wang, Jinqiang Cui, Xinlei Chen, et al\.2025\.Urbanvideo\-bench: Benchmarking vision\-language models on embodied intelligence with video data in urban spaces\. In*Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*\. 32400–32423\.[doi:10\.18653/v1/2025\.acl\-long\.1558](https://doi.org/10.18653/v1/2025.acl-long.1558)
- Zheng et al\.\(2022\)Kaizhi Zheng, Xiaotong Chen, Odest Chadwicke Jenkins, and Xin Wang\. 2022\.Vlmbench: A compositional benchmark for vision\-and\-language manipulation\.*Advances in Neural Information Processing Systems*35 \(2022\), 665–678\.[https://dl\.acm\.org/doi/10\.5555/3600270\.3600318](https://dl.acm.org/doi/10.5555/3600270.3600318)
- Zhou et al\.\(2024\)Qinhong Zhou, Sunli Chen, Yisong Wang, Haozhe Xu, Weihua Du, Hongxin Zhang, Yilun Du, Joshua B Tenenbaum, and Chuang Gan\. 2024\.Hazard challenge: Embodied decision making in dynamically changing environments\. In*International Conference on Learning Representations*\.[https://research\.ibm\.com/publications/hazard\-challenge\-embodied\-decision\-making\-in\-dynamically\-changing\-environments](https://research.ibm.com/publications/hazard-challenge-embodied-decision-making-in-dynamically-changing-environments)
- Zhu et al\.\(2023\)Wanrong Zhu, Jack Hessel, Anas Awadalla, Samir Yitzhak Gadre, Jesse Dodge, Alex Fang, Youngjae Yu, Ludwig Schmidt, William Yang Wang, and Yejin Choi\. 2023\.Multimodal c4: An open, billion\-scale corpus of images interleaved with text\.*Advances in Neural Information Processing Systems*36 \(2023\), 8958–8974\.[https://dl\.acm\.org/doi/10\.5555/3666122\.3666515](https://dl.acm.org/doi/10.5555/3666122.3666515)

Similar Articles

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes

Hugging Face Daily Papers

SpatialAct is a new simulator-grounded benchmark that probes whether VLM agents can perform coherent spatial reasoning and translate it into actions in 3D environments across multi-turn feedback settings. Experiments reveal a significant reasoning-to-action gap, with current VLMs struggling to maintain spatial beliefs and produce reliable actions despite performing well on isolated reasoning tasks.