ABot-N1: Toward a General Visual Language Navigation Foundation Model
Summary
ABot-N1 is a visual language navigation foundation model that decouples cognition from control using a slow-fast architecture with dual visual-language signals, achieving new state-of-the-art results across indoor and outdoor navigation tasks.
View Cached Full Text
Cached at: 07/14/26, 04:13 AM
Paper page - ABot-N1: Toward a General Visual Language Navigation Foundation Model
Source: https://huggingface.co/papers/2607.10383 Published on Jul 11
#2 Paper of the day Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
VisualLanguageNavigationfoundationmodelsaimtounifydeepreasoningforgroundedspatialdecisionswithbroadversatilityfordiverseembodiedtasks.Currentapproachestypicallyachievethisintegrationviamonolithicpoliciesthatmapobservationsdirectlytoactions,yettheyoftensufferfromcoordinatedriftandpoorhandlingoflong-tailsemantics.Furthermore,theseblack-boxmappingslackinterpretability,hinderingthesimultaneousachievementofgenerality,robustness,andtransparency.WepresentABot-N1,asteptowardageneralVisualLanguageNavigationfoundationmodel,thataddressesthesechallengesbydecouplingcognitionfromcontrolviaaslow-fastarchitectureguidedbydualvisual-languagesignals.Morespecifically,aslowvision-languagereasonerperformsexplicitChain-of-Thoughtreasoningwhileproducingapixelgoal.Thiscompactsetofimage-spaceanchorpointsservesasauniversalinterfacefordiversetasks,includingpoint-goal,object-goal,poi-goal,instruction-following,andperson-following.Subsequently,afastactionexpertleveragesboththetextualcuesandthepixelguidancetogeneratecontinuouswaypointsatthenativecontrolfrequency.Bybridginghigh-levelintentsandlow-levelcontrolthroughpixel-groundedanchorspairedwithexplicitlinguistictraces,ourapproachensuresrobust,generalizable,andinterpretablenavigationacrosssimulationandreal-worldbenchmarks.ABot-N1establishesnewstate-of-the-artrecords,deliveringmassivegainsspecificallyinurban-scalenavigation:boostingPOIarrivalby35.0%(to77.3%)andachieving95.4%/92.9%SRincomplexindoorandoutdoorscenes.Italsomaintainssuperiorrobustnessacrossobject-reaching,person-following,andinstruction-followingtasks.NewPoint-Goal/POI-Goalbenchmarksarereleasedasopensourcetoadvancethefieldofurban-scalenavigation.
View arXiv pageView PDFProject pageAdd to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.10383 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.10383 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.10383 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
ABot-M0.5: Unified Mobility-and-Manipulation World Action Model
ABot-M0.5 is a new World Action Model for mobile manipulation that improves performance through temporal granularity alignment, action space disentanglement, and train-test consistency, achieving state-of-the-art results on long-horizon and fine-grained manipulation benchmarks.
PlatonicNav: Unveiling Semantic Correspondence in Navigation with Platonic Topological Maps
PlatonicNav introduces a training-free framework for embodied navigation that uses vision-only semantic maps and blind matching to ground language goals, achieving generalization across tasks and embodiments without explicit cross-modal training.
@AdinaYakup: LingBot Vision A self-supervised vision backbone family for dense spatial perception from Ant Group @robbyant_brain - A…
LingBot Vision, a self-supervised vision backbone family from Ant Group, uses masked boundary modeling to achieve state-of-the-art performance on dense spatial perception tasks, beating the larger DINOv3 model on NYU-Depth v2.
Qwen-RobotNav Technical Report: A Scalable Navigation Model Designed for an Agentic Navigation System
Qwen-RobotNav is a scalable navigation model with a parameterized interface enabling dynamic task modes and observation parameters, achieving state-of-the-art performance through multi-task training and zero-shot generalization to real-world robotics.
Path-level Hindsight Instructions for Semantic Exploration in Vision-Language Navigation
Introduces Φ-Nav, a unified on-policy framework that uses hindsight reasoning to synthetically generate path-level instructions from exploratory trajectories, bridging the semantic supervision gap in Vision-Language Navigation and achieving competitive results on R2R-CE and RxR-CE benchmarks with fewer expert demonstrations.