Tag
This paper proposes a collaborative memory framework for multi-agent vision-language model systems to address distributed perception and improve shared visual context and reasoning consistency.
This paper introduces GPT-Policy, a framework for in-context robot learning using vision-language models, enabling robots to learn from demonstrations without gradient updates. It evaluates the framework in real-robot trials, showing improved task completion.
The paper introduces MineAmongUs, a 3D multimodal Among Us environment, and the ARIA harness to study deception in VLM agents, finding that non-verbal actions are key to successful deception in social interactions.
This paper empirically studies how VLM agents with persistent spatial memory fail when memory becomes stale, using a dynamic FrozenLake testbed. It finds that trusting stale memory can more than double death rates, and that read-time auditing helps but does not fully close the gap.
This paper empirically studies spatial memory staleness in vision-language-model agents, finding that models often ignore contradictory visual evidence and that trusting stale memory can increase safety risks. The authors propose auditing mechanisms but show that visual grounding under memory-observation conflicts remains a major open challenge.
SceneActBench is a benchmark for evaluating VLM agents on acting in complete multi-object 3D scenes, using task-specific geometric metrics across five tasks.
GROW proposes a novel reinforcement learning framework that adapts GRPO to multi-turn VLM agent tasks by decomposing trajectories into state-action pairs and computing advantages between them, achieving state-of-the-art performance on over 800 Minecraft tasks.
AtlasVA is a teacher-free visual skill memory framework for vision-language model agents that uses spatial heatmaps, visual exemplars, and symbolic text skills to improve spatial decision-making in long-horizon tasks, outperforming baselines on several benchmarks.