Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation

Hugging Face Daily Papers Papers

Summary

TAMP-Nav is a unified framework that enhances embodied navigation by aligning vision-language models with 2D visual prompting, selective reasoning with compressed memory, and policy optimization, achieving state-of-the-art performance with high runtime and sample efficiency.

Although Large Vision-Language Models (VLMs) have significantly advanced embodied navigation, their direct deployment remains challenging, as existing methods often force VLMs into unnatural action spaces that misalign with their 2D pre-training priors, compounded by rigid reasoning schedules and inefficient memory management. To overcome these limitations, we propose TAMP-Nav, a unified framework for efficient embodied navigation. First, we introduce a Pixel-to-3D Action Formulation (Point) that reformulates navigation into 2D visual prompting. Specifically, the VLM merely selects 2D pixels, which are then projected into 3D coordinates for a low-level SLAM controller. This design naturally aligns embodied execution with the VLM's inherent 2D visual capabilities. Second, we propose an integrated Selective Reasoning and Anchor-Trajectory Memory mechanism (Think and Memorize), which dynamically triggers Chain-of-Thought and retains high-fidelity memory only at critical nodes, compressing redundant trajectories into lightweight Space-Time Indicators, thereby preserving critical historical information and enhancing spatio-temporal perception. Finally, we design an efficient Two-Level Alignment Paradigm (Align) via Group Relative Policy Optimization (GRPO). By superimposing global outcome rewards with fine-grained process rewards, this dense supervision tightly aligns the agent's cognitive planning with physical environmental feedback, endowing the model with adaptive reasoning capabilities. Experiments demonstrate that TAMP-Nav achieves state-of-the-art performance (e.g., 66.2% SR on R2R-CE) with high runtime and sample efficiency (requiring only 90k training trajectories).
Original Article
View Cached Full Text

Cached at: 08/19/26, 03:56 AM

Paper page - Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation

Source: https://huggingface.co/papers/2608.17512 Authors:

,

,

,

,

,

,

,

,

,

,

Abstract

TAMP-Nav improves embodied navigation by aligning vision-language models with 2D visual prompting, selective reasoning with compressed memory, and dense policy optimization.

AlthoughLarge Vision-Language Models(VLMs) have significantly advanced embodied navigation, their direct deployment remains challenging, as existing methods often force VLMs into unnatural action spaces that misalign with their 2D pre-training priors, compounded by rigid reasoning schedules and inefficient memory management. To overcome these limitations, we propose TAMP-Nav, a unified framework for efficient embodied navigation. First, we introduce aPixel-to-3D Action Formulation(Point) that reformulates navigation into2D visual prompting. Specifically, the VLM merely selects 2D pixels, which are then projected into 3D coordinates for a low-levelSLAM controller. This design naturally aligns embodied execution with the VLM’s inherent 2D visual capabilities. Second, we propose an integrated Selective Reasoning andAnchor-Trajectory Memorymechanism (Think and Memorize), which dynamically triggersChain-of-Thoughtand retains high-fidelity memory only at critical nodes, compressing redundant trajectories into lightweightSpace-Time Indicators, thereby preserving critical historical information and enhancing spatio-temporal perception. Finally, we design an efficient Two-Level Alignment Paradigm (Align) viaGroup Relative Policy Optimization(GRPO). By superimposingglobal outcome rewardswith fine-grainedprocess rewards, this dense supervision tightly aligns the agent’s cognitive planning with physical environmental feedback, endowing the model with adaptive reasoning capabilities. Experiments demonstrate that TAMP-Nav achieves state-of-the-art performance (e.g., 66.2% SR on R2R-CE) with high runtime and sample efficiency (requiring only 90k training trajectories).

View arXiv pageView PDFProject pageGitHubAdd to collection

Get this paper in your agent:

hf papers read 2608\.17512

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2608.17512 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2608.17512 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2608.17512 in a Space README.md to link it from this page.

Collections including this paper1

Similar Articles

NavHarness: Towards Lifelong Embodied Navigation

Hugging Face Daily Papers

The paper introduces NavHarness, a training-free embodied navigation harness that integrates memory processing into the navigation loop, enabling agents to carry experience across sessions for lifelong navigation. On GOAT-Bench, it improves success rate by up to 22.6 points over context-only sessions and achieves state-of-the-art results with GPT-6/Astra.

Beyond the Remembered World: Predictive 4D Belief for Persistent Navigation in Evolving Worlds

Hugging Face Daily Papers

The paper introduces EvolvingNav, a predictive 4D belief framework for embodied agents that tracks targets in evolving environments, forecasts their locations at inspection time, and revises beliefs under limited visibility, along with the EvoWorld-Bench benchmark covering 54 scenes and 803,680 tasks, demonstrating improved navigation success in simulation and real-robot experiments.