ABot-N1: Toward a General Visual Language Navigation Foundation Model
Summary
ABot-N1 is a visual language navigation foundation model that decouples cognition from control using a slow-fast architecture with dual visual-language signals, achieving new state-of-the-art results across indoor and outdoor navigation tasks.
View Cached Full Text
Cached at: 07/14/26, 04:13 AM
Paper page - ABot-N1: Toward a General Visual Language Navigation Foundation Model
Source: https://huggingface.co/papers/2607.10383 Published on Jul 11
#2 Paper of the day Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
VisualLanguageNavigationfoundationmodelsaimtounifydeepreasoningforgroundedspatialdecisionswithbroadversatilityfordiverseembodiedtasks.Currentapproachestypicallyachievethisintegrationviamonolithicpoliciesthatmapobservationsdirectlytoactions,yettheyoftensufferfromcoordinatedriftandpoorhandlingoflong-tailsemantics.Furthermore,theseblack-boxmappingslackinterpretability,hinderingthesimultaneousachievementofgenerality,robustness,andtransparency.WepresentABot-N1,asteptowardageneralVisualLanguageNavigationfoundationmodel,thataddressesthesechallengesbydecouplingcognitionfromcontrolviaaslow-fastarchitectureguidedbydualvisual-languagesignals.Morespecifically,aslowvision-languagereasonerperformsexplicitChain-of-Thoughtreasoningwhileproducingapixelgoal.Thiscompactsetofimage-spaceanchorpointsservesasauniversalinterfacefordiversetasks,includingpoint-goal,object-goal,poi-goal,instruction-following,andperson-following.Subsequently,afastactionexpertleveragesboththetextualcuesandthepixelguidancetogeneratecontinuouswaypointsatthenativecontrolfrequency.Bybridginghigh-levelintentsandlow-levelcontrolthroughpixel-groundedanchorspairedwithexplicitlinguistictraces,ourapproachensuresrobust,generalizable,andinterpretablenavigationacrosssimulationandreal-worldbenchmarks.ABot-N1establishesnewstate-of-the-artrecords,deliveringmassivegainsspecificallyinurban-scalenavigation:boostingPOIarrivalby35.0%(to77.3%)andachieving95.4%/92.9%SRincomplexindoorandoutdoorscenes.Italsomaintainssuperiorrobustnessacrossobject-reaching,person-following,andinstruction-followingtasks.NewPoint-Goal/POI-Goalbenchmarksarereleasedasopensourcetoadvancethefieldofurban-scalenavigation.
View arXiv pageView PDFProject pageAdd to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.10383 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.10383 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.10383 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation
LightNav-0 is a compact generalist navigation model that leverages pretrained vision-language models' spatial intelligence to achieve state-of-the-art embodied navigation across diverse tasks and robot embodiments.
Scaffolding Foundation Models into Physical-World Agents Pushes the Frontier of Long-Horizon Navigation
NavMCP is an agentic scaffolding framework that integrates vision-language models with navigation foundation models to enable persistent long-horizon physical-world exploration, achieving state-of-the-art results on benchmarks and real-robot tasks.
Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation
TAMP-Nav is a unified framework that enhances embodied navigation by aligning vision-language models with 2D visual prompting, selective reasoning with compressed memory, and policy optimization, achieving state-of-the-art performance with high runtime and sample efficiency.
ABot-M0.5: Unified Mobility-and-Manipulation World Action Model
ABot-M0.5 is a new World Action Model for mobile manipulation that improves performance through temporal granularity alignment, action space disentanglement, and train-test consistency, achieving state-of-the-art results on long-horizon and fine-grained manipulation benchmarks.
PlatonicNav: Unveiling Semantic Correspondence in Navigation with Platonic Topological Maps
PlatonicNav introduces a training-free framework for embodied navigation that uses vision-only semantic maps and blind matching to ground language goals, achieving generalization across tasks and embodiments without explicit cross-modal training.