Tag
VABench introduces a benchmark to evaluate embodied spatial intelligence in models by testing their ability to observe, reason, and act through visual demonstrations and active perception. It shows that active camera control improves task success, but no model completes long-horizon episodes.
PhysCaP is a physics-informed code-generation agent that actively explores objects to infer hidden physical properties for efficient robotic manipulation.
This paper proposes an active-perception framework for embodied target disambiguation, using vision-language models to decide based on accumulated visual evidence and interaction information. Real-robot experiments demonstrate its effectiveness in combining physical observation with user-intent clarification.
This paper proposes a mathematical formulation of slow thinking and active perception, introducing a theory called 'active lifting' that derives design, training, and inference processes for slow thinking large language models.
The paper introduces a certified sensing clock for world models that provides a proactive, analytically-derived deadline for when an agent must re-sense, based on the model's validity horizon. It proposes a drift-aware deployment method and demonstrates its effectiveness on a frozen 3D VN-JEPA model, reducing tail violations compared to baseline schedulers.
Introduces OmniAgent, an omni-modal agent that uses an iterative Observation-Thought-Action cycle with active perception to achieve superior long video understanding, outperforming larger models like Qwen2.5-VL-72B on benchmarks.
Co-GLANCE is a real-time onboard perception and decision-making system for heterogeneous robot teams that distills vision-language model capabilities into efficient models and uses conformal prediction with selective abstention to quantify and resolve perceptual uncertainty, outperforming cloud-based VLM baselines by 25-36% while achieving 350x lower latency.
ActiveMimic is a pretraining framework that recovers camera and wrist trajectories from egocentric human video to model active perception as a viewpoint action, enabling robot pretraining that matches the performance of models trained directly on robot data.