Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence
Summary
This paper introduces SIS-Bench, a benchmark for evaluating self-awareness and spatial cognition in UAV embodied intelligence using multimodal large language models, and explores motion-aware representations to improve performance.
View Cached Full Text
Cached at: 07/16/26, 05:44 PM
Paper page - Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence
Source: https://huggingface.co/papers/2607.12477
Abstract
AutonomousUAVsystemsincreasinglyrelyonmultimodallargelanguagemodels(MLLMs)tooperateincomplexreal-worldenvironments.Suchembodiedscenariosrequirenotonlyunderstandingthesurroundingspacebutalsomaintainingacoherentrepresentationoftheagentitself.However,existingUAV-orientedapproachesandbenchmarksremainlargelyenvironment-centric,primarilyfocusingonspatialunderstandingtasks,withtheagent’sself-awarenessremainingimplicit.Toaddressthisgap,weintroduceSIS-Bench,abenchmarkforevaluatingembodiedspatialintelligenceinUAVscenariosunderaunifiedself-in-spaceformulation.SIS-Benchorganizesevaluationalongtwocomplementarydimensions,spaceandself,andathree-levelhierarchyofperception,memory,andreasoning.Itcontains4,856question--answerpairsacross13tasksderivedfrom1,646real-worldUAVvideosthroughatask-conditionedconstructionpipelinewithexpertverification.ExtensiveevaluationsrevealthatcurrentMLLMsexhibitfundamentallimitationsinmodelingdynamicandagent-centeredprocesses.Inparticular,weobserveaclearimbalancebetweenspatialcognitionandself-awareness,aswellasaprogressiveperformancedegradationacrosscognitivelevels.Motivatedbythesefindings,wefurtherexploreamotion-awarerepresentationthatincorporatesself-relateddynamicsthroughopticalflowandvisualfeaturefusion.Experimentalresultsshowthatmodelingagentmotionconsistentlyimprovesperceptionandmemoryperformance,notonlyinspatialcognitionbutalsoinself-awareness,andgeneralizestodownstreamUAVdecision-makingtasks.Ourresultshighlighttheimportanceofself-awarenessforadvancingembodiedspatialintelligence,andprovidebothanewbenchmarkandempiricalevidenceformotion-awareself-in-spacemodeling.
View arXiv pageView PDFProject pageGitHub4Add to collection
Get this paper in your agent:
hf papers read 2607\.12477
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper1
#### choucsan/SIS-Motion Updated1 day ago • 88 • 2
Datasets citing this paper3
#### choucsan/SIS-Bench Viewer• Updatedabout 10 hours ago • 4.86k • 591 • 2 #### choucsan/OpenUAV-QA Preview• Updated1 day ago • 372 • 2 #### choucsan/SIS-Motion-54K Viewer• Updated1 day ago • 54.3k • 167 • 4
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.12477 in a Space README.md to link it from this page.
Collections including this paper1
Similar Articles
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop
Introduces ESI-BENCH, a comprehensive benchmark for embodied spatial intelligence built on OmniGibson, covering 10 task categories and 29 subcategories. Experiments show active exploration substantially outperforms passive approaches, with failures mainly due to action blindness rather than perception, revealing a metacognitive gap in models compared to humans.
Embodied-BenchClaw: An Autonomous Multi-Agent System for Embodied Spatial Intelligence Benchmark Construction
This paper proposes Embodied-BenchClaw, an autonomous multi-agent system that automatically constructs embodied spatial intelligence benchmarks from user intent through a five-stage pipeline with process quality control and an extensible Skill Library.
SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks
SpatialWorld is a unified benchmark for evaluating interactive spatial reasoning in multimodal agents across diverse real-world tasks, revealing that even the strongest models achieve low task success rates.
OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs
OVO-S-Bench introduces a comprehensive human-annotated benchmark of 1,680 questions across 348 videos to evaluate streaming spatial intelligence in multimodal LLMs, revealing that even the best model (Gemini-3.1-Pro) trails human experts by 27 points. The benchmark exposes key limitations including allocentric mapping as a major bottleneck and chain-of-thought reasoning amplifying spatial errors.
UESF-Bench: Benchmarking and Probing for Unified Embodied Seeking and Following
This paper introduces UESF-Bench, a large-scale benchmark for unified embodied human seeking and following, and proposes SeekFollow-VLA, a vision-language-action framework that handles semantic-guided exploration and reliable behavior switching between searching and following.