VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control
Summary
VABench introduces a benchmark to evaluate embodied spatial intelligence in models by testing their ability to observe, reason, and act through visual demonstrations and active perception. It shows that active camera control improves task success, but no model completes long-horizon episodes.
View Cached Full Text
Cached at: 09/18/26, 03:01 AM
Paper page - VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control
Source: https://huggingface.co/papers/2609.19554
Abstract
Spatialintelligencerequiresmorethandescribingobjectlocations.Underincompleteobservation,modelsmustidentifyandacquiremissingevidence,interpretitinacommonspatialframe,andactonit.WeintroduceVA-Benchtoevaluatethecompleteobserve-reason-act-reviseloop.General-purposeMLLMslearnproceduralcontextfromRGB-onlydemonstrations,activelyselectcameraviewpoints,issuemetricCartesiancommands,andrevisethemfromexecutionfeedback.Modelsreceivenoprivilegedobjectposes,oracletrajectories,orlearnedactionheads.Afixedmodel-agnosticcontrollerexecutesonlymodel-specifiedtargets.VA-Benchcontains14basetaskfamilies(11single-armandthreedual-arm),sevenheld-outgeometry/layoutvariants,andalong-horizonfive-objectcompositiontrack.Weevaluate12primarymodelconditionsinthreeindependentrunsoverthesame20physicallyverifiedseedsperbasetask,reportingterminalsuccess,ninetrajectory-levelbehavioraldiagnostics,andsubtaskprogress.First,thebest-performingmodelscores100.0%ontargetlocalizationand78.9%onspatialrelationsintheannotatedrun.Itsthree-runmacro-averagetasksuccessisonly53.93+/-3.17%.Second,activecameracontrolsignificantlyimprovestasksuccessoverpassivemulti-viewobservation.Inonematchedcomparison,successrisesfrom27.86%to57.50%.Third,held-outgeometrictransfercanreducetasksuccessbyover30percentagepoints.Nomodelcompletesastrictlong-horizonepisode,despitesubstantialpartialprogress.VA-Benchthustestswhethergeneral-purposeMLLMscanturnvisualdemonstrationsandactivelyacquiredevidenceintosuccessfulembodiedaction.
View arXiv pageView PDFGitHub3Add to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.19554 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.19554 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.19554 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop
Introduces ESI-BENCH, a comprehensive benchmark for embodied spatial intelligence built on OmniGibson, covering 10 task categories and 29 subcategories. Experiments show active exploration substantially outperforms passive approaches, with failures mainly due to action blindness rather than perception, revealing a metacognitive gap in models compared to humans.
Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence
This paper introduces SIS-Bench, a benchmark for evaluating self-awareness and spatial cognition in UAV embodied intelligence using multimodal large language models, and explores motion-aware representations to improve performance.
Embodied-BenchClaw: An Autonomous Multi-Agent System for Embodied Spatial Intelligence Benchmark Construction
This paper proposes Embodied-BenchClaw, an autonomous multi-agent system that automatically constructs embodied spatial intelligence benchmarks from user intent through a five-stage pipeline with process quality control and an extensible Skill Library.
SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence
Introduces SVI-Bench, a large-scale benchmark for strategic video intelligence using team sports, designed to evaluate models on dynamic scene understanding, causal reasoning, strategic simulation, and agentic synthesis. The benchmark reveals a capability cliff where models perform well on perceptual tasks but sharply degrade on higher-level strategic reasoning.
SceneActBench: Can Agents Act on the 3D Scenes They See?
SceneActBench is a benchmark for evaluating VLM agents on acting in complete multi-object 3D scenes, using task-specific geometric metrics across five tasks.