HarnessVLN: Unifying Training-Free Embodied Navigation through an Agent Harness

Hugging Face Daily Papers Papers

Summary

HarnessVLN is a zero-shot, training-free framework for embodied navigation that unifies perception, retrieval, grounding, navigation, recovery, and termination through a unified tool interface, achieving state-of-the-art results on benchmarks like R2R and RxR.

Embodied navigation requires agents to interpret visual observations, accumulate spatial knowledge, and execute actions to follow instructions or locate objects. Training-based methods face generalization challenges, while training-free methods exploit multimodal large language models (MLLMs) but often lack mechanisms to reconcile proposed actions with spatial evidence, task progress, and execution failures. We present HarnessVLN, a zero-shot, training-free framework whose Agent Harness coordinates perception, retrieval, grounding, navigation, recovery, and termination through a unified tool interface. The Harness validates planner proposals against spatial evidence, geometric feasibility, and subgoal consistency, incorporating structured tool feedback into subsequent decisions. Hierarchical event memory tracks task progress and execution history, while a persistent Spatiotemporal Graph maintains reusable spatial evidence and failure annotations for verification and recovery. A replaceable Navigation Executor converts validated targets into executable motions, allowing the same Harness protocol to support instruction-following and object-goal navigation. HarnessVLN achieves success rates of 60.8%, 53.9%, 76.0%, and 59.3% on R2R, RxR, HM3D-v2, and HM3D-OVON, respectively, surpassing prior training-free SOTA results. Humanoid deployment further demonstrates its applicability to both tasks in real-world environments. The project page is: https://harnessvln.netlify.app/.
Original Article
View Cached Full Text

Cached at: 09/16/26, 10:47 AM

Paper page - HarnessVLN: Unifying Training-Free Embodied Navigation through an Agent Harness

Source: https://huggingface.co/papers/2609.15195

Abstract

Embodiednavigationrequiresagentstointerpretvisualobservations,accumulatespatialknowledge,andexecuteactionstofollowinstructionsorlocateobjects.Training-basedmethodsfacegeneralizationchallenges,whiletraining-freemethodsexploitmultimodallargelanguagemodels(MLLMs)butoftenlackmechanismstoreconcileproposedactionswithspatialevidence,taskprogress,andexecutionfailures.WepresentHarnessVLN,azero-shot,training-freeframeworkwhoseAgentHarnesscoordinatesperception,retrieval,grounding,navigation,recovery,andterminationthroughaunifiedtoolinterface.TheHarnessvalidatesplannerproposalsagainstspatialevidence,geometricfeasibility,andsubgoalconsistency,incorporatingstructuredtoolfeedbackintosubsequentdecisions.Hierarchicaleventmemorytrackstaskprogressandexecutionhistory,whileapersistentSpatiotemporalGraphmaintainsreusablespatialevidenceandfailureannotationsforverificationandrecovery.AreplaceableNavigationExecutorconvertsvalidatedtargetsintoexecutablemotions,allowingthesameHarnessprotocoltosupportinstruction-followingandobject-goalnavigation.HarnessVLNachievessuccessratesof60.8%,53.9%,76.0%,and59.3%onR2R,RxR,HM3D-v2,andHM3D-OVON,respectively,surpassingpriortraining-freeSOTAresults.Humanoiddeploymentfurtherdemonstratesitsapplicabilitytobothtasksinreal-worldenvironments.Theprojectpageis:https://harnessvln.netlify.app/.

View arXiv pageView PDFProject pageAdd to collection

Get this paper in your agent:

hf papers read 2609\.15195

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.15195 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.15195 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.15195 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Show-Harness: Just a VLM Agent Can Play Robots

Hugging Face Daily Papers

Show-Harness is a method that enables vision-language models to control robots through discrete semantic actions, allowing zero-shot deployment and efficient fine-tuning across different robots and GUIs.

HarnessBridge: Learnable Bidirectional Controller for LLM Agent Harness

Hugging Face Daily Papers

Introduces HarnessBridge, a learnable bidirectional controller that parameterizes the agent-environment interface for LLM agents, achieving performance comparable to specialized harnesses with reduced computational overhead on Terminal-Bench and SWE-bench.