Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation
Summary
TAMP-Nav is a unified framework that enhances embodied navigation by aligning vision-language models with 2D visual prompting, selective reasoning with compressed memory, and policy optimization, achieving state-of-the-art performance with high runtime and sample efficiency.
View Cached Full Text
Cached at: 08/19/26, 03:56 AM
Paper page - Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation
Source: https://huggingface.co/papers/2608.17512 Authors:
,
,
,
,
,
,
,
,
,
,
Abstract
TAMP-Nav improves embodied navigation by aligning vision-language models with 2D visual prompting, selective reasoning with compressed memory, and dense policy optimization.
AlthoughLarge Vision-Language Models(VLMs) have significantly advanced embodied navigation, their direct deployment remains challenging, as existing methods often force VLMs into unnatural action spaces that misalign with their 2D pre-training priors, compounded by rigid reasoning schedules and inefficient memory management. To overcome these limitations, we propose TAMP-Nav, a unified framework for efficient embodied navigation. First, we introduce aPixel-to-3D Action Formulation(Point) that reformulates navigation into2D visual prompting. Specifically, the VLM merely selects 2D pixels, which are then projected into 3D coordinates for a low-levelSLAM controller. This design naturally aligns embodied execution with the VLM’s inherent 2D visual capabilities. Second, we propose an integrated Selective Reasoning andAnchor-Trajectory Memorymechanism (Think and Memorize), which dynamically triggersChain-of-Thoughtand retains high-fidelity memory only at critical nodes, compressing redundant trajectories into lightweightSpace-Time Indicators, thereby preserving critical historical information and enhancing spatio-temporal perception. Finally, we design an efficient Two-Level Alignment Paradigm (Align) viaGroup Relative Policy Optimization(GRPO). By superimposingglobal outcome rewardswith fine-grainedprocess rewards, this dense supervision tightly aligns the agent’s cognitive planning with physical environmental feedback, endowing the model with adaptive reasoning capabilities. Experiments demonstrate that TAMP-Nav achieves state-of-the-art performance (e.g., 66.2% SR on R2R-CE) with high runtime and sample efficiency (requiring only 90k training trajectories).
View arXiv pageView PDFProject pageGitHubAdd to collection
Get this paper in your agent:
hf papers read 2608\.17512
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.17512 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2608.17512 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.17512 in a Space README.md to link it from this page.
Collections including this paper1
Similar Articles
NavHarness: Towards Lifelong Embodied Navigation
The paper introduces NavHarness, a training-free embodied navigation harness that integrates memory processing into the navigation loop, enabling agents to carry experience across sessions for lifelong navigation. On GOAT-Bench, it improves success rate by up to 22.6 points over context-only sessions and achieves state-of-the-art results with GPT-6/Astra.
From Passive Retrieval to Active Memory Navigation: Learning to Use Memory as a Structured Action Space
NapMem is a framework that treats long-term user memory as a structured action space rather than passive retrieval, using a multi-granularity memory pyramid and reinforcement learning to train agents to navigate memory. Experiments show competitive performance on memory-intensive tasks.
Beyond the Remembered World: Predictive 4D Belief for Persistent Navigation in Evolving Worlds
The paper introduces EvolvingNav, a predictive 4D belief framework for embodied agents that tracks targets in evolving environments, forecasts their locations at inspection time, and revises beliefs under limited visibility, along with the EvoWorld-Bench benchmark covering 54 scenes and 803,680 tasks, demonstrating improved navigation success in simulation and real-robot experiments.
ABot-N1: Toward a General Visual Language Navigation Foundation Model
ABot-N1 is a visual language navigation foundation model that decouples cognition from control using a slow-fast architecture with dual visual-language signals, achieving new state-of-the-art results across indoor and outdoor navigation tasks.
LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation
LightNav-0 is a compact generalist navigation model that leverages pretrained vision-language models' spatial intelligence to achieve state-of-the-art embodied navigation across diverse tasks and robot embodiments.