Grounded Action Model: 3D Grounding as a Foundation for Robotics
Summary
This paper proposes Grounded Action Models (GAMs), a new paradigm for robot foundation models that integrates 3D grounding, achieving state-of-the-art performance on manipulation tasks.
View Cached Full Text
Cached at: 09/22/26, 07:27 AM
Paper page - Grounded Action Model: 3D Grounding as a Foundation for Robotics
Source: https://huggingface.co/papers/2609.23863 Published on Sep 20
·
Submitted byhttps://huggingface.co/Jiafei1224
Duanon Sep 22
Abstract
Manipulationpoliciesmustknowwhichobjectsmatterandwheretheyare,yetthepretrainedbackbonesthatcurrentrobotfoundationmodelsbuildon,fromlanguageinvision-language-actionmodels(VLAs)tovideogenerationinworld-actionmodels(WAMs),donotdirectlyrequirethismetricgrounding,leavingittobelearnedimplicitlyfromrobotdemonstrations.WeproposeGroundedActionModels(GAMs),anewparadigmofrobotfoundationmodelsbuiltwith3Dgrounding.GAMcanbeconditionedusinglanguage,points,orboxprompts,whicharefirsttransformedintoasharedobject-centricrepresentationoftheselectedobjects.Thisrepresentationcapturestarget-focusedvisualfeaturesandmetricobjectgeometry,whichismixedwithrobotstatehistorythroughamulti-streamtransformertopredictactionchunks.AlthoughGAMscanberunautonomously,theycanalsoserveasalow-levelcontrollerthatahigh-levelplannercontrolsusingitsvariousinputmodalities,allowingforlong-horizonandmemory-dependentmanipulation.OnRoboTwin2.0,GAMachievesanaveragesuccessrateof55.3%across50tasks(vs.52.0%forSpatialForcing),including47.6%underscenerandomization(vs.30.4%forAbot-M0),withitsactionpolicytrainedonlyonclean-scenedemonstrations.OnLIBERO-PRO,itachievesastate-of-the-artaveragesuccessrateof61%(vs.53%forπ_{0.5})across16perturbationsettings,withthelargestgainswhentargetsarerelocatedornewlydesignated.Ontworealrobots,GAMretains17/20successesundervisualshiftonabimanualYAMversus4/20forπ_{0.5},whileitscompositionwithaMolmo2planneronaFrankaachieves64.7%IDand49.8%OODstepcompletiononlong-horizonandmemory-dependenttasks.
View arXiv pageView PDFProject pageGitHub0Add to collection
Get this paper in your agent:
hf papers read 2609\.23863
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.23863 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.23863 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.23863 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Geometric Action Model for Robot Policy Learning
The Geometric Action Model (GAM) repurposes a pretrained geometric foundation model (GFM) as a unified backbone for language-conditioned robot manipulation, achieving higher accuracy, robustness, and efficiency than existing foundation-model-scale baselines across simulation and real-world benchmarks.
GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation
This paper introduces GE-Act 2.0, a world-action model pretrained from scratch to enable scalable zero-shot robotic manipulation with improved success rates across diverse tasks and conditions.
Cortex 2.0: Grounding World Models in Real-World Industrial Deployment
Cortex 2.0 introduces a plan-and-act control framework that uses visual latent space trajectory generation to enable reliable long-horizon robotic manipulation in complex industrial environments, outperforming reactive Vision-Language-Action models.
Revisiting Articulated Parts Perception in Robot Manipulation
This paper introduces Geometric Primary Structure (GPS), a new representation for articulated parts perception in robot manipulation, enabling efficient VR-based annotation and achieving a 73% success rate without fine-tuning.
ABot-M0.5: Unified Mobility-and-Manipulation World Action Model
ABot-M0.5 is a new World Action Model for mobile manipulation that improves performance through temporal granularity alignment, action space disentanglement, and train-test consistency, achieving state-of-the-art results on long-horizon and fine-grained manipulation benchmarks.