Tag
This paper proposes Grounded Action Models (GAMs), a new paradigm for robot foundation models that integrates 3D grounding, achieving state-of-the-art performance on manipulation tasks.
This paper introduces SR-REAL, a unified framework for spatial vision-language models that combines linguistic deduction and 3D geometric reasoning via reinforcement learning, enabling robust multi-step spatial reasoning across diverse tasks.