Beyond the Remembered World: Predictive 4D Belief for Persistent Navigation in Evolving Worlds
Summary
The paper introduces EvolvingNav, a predictive 4D belief framework for embodied agents that tracks targets in evolving environments, forecasts their locations at inspection time, and revises beliefs under limited visibility, along with the EvoWorld-Bench benchmark covering 54 scenes and 803,680 tasks, demonstrating improved navigation success in simulation and real-robot experiments.
View Cached Full Text
Cached at: 10/01/26, 12:23 PM
Paper page - Beyond the Remembered World: Predictive 4D Belief for Persistent Navigation in Evolving Worlds
Source: https://huggingface.co/papers/2609.39166
Abstract
Persistentspatialmemoryenablesembodiedagentstonavigatefamiliarenvironmentsacrossrepeatedvisits.However,targetsmaymovewhileunobserved,includingduringnavigation,makingrememberedlocationsunreliablebythetimeanagentarrives.Despiteadvancesinmemoryretrievalandstateprediction,accountingforcontinuedhiddenworldevolutionandrevisingbeliefsunderlimitedvisibilityremainchallenging.WestudyEvolving-WorldNavigation,whereagentsinfertargetlocationsfromintermittentobservations,predicttheirstatesatinspectiontime,andrevisebeliefsusingvisualevidence.WeproposeEvolvingNav,whichconstructsatime-indexedbelieffromtimestamped3Dobjecthistoriesthroughastructuredpersistence-relocationmodel.Thebeliefdistinguishespersistenceatthelastobservedlocationfromrelocationtoalternativelocationsandretainsprobabilitymassoutsidetheknowncandidateset.Anevent-drivenfilterpropagatesthecurrentbeliefastimeelapses,forecaststargetoccupancyatcandidateinspectiontimes,andincorporatesnewRGB-Devidence.Negativeobservationsdownweightlocationhypothesesaccordingtocalibrated,visibility-conditioneddetectionprobabilities,whileevidencetrackingpreventsrepeateduseofthesameobservations.Afrozen,zero-shotvision-languagecontrollerusestheupdatedbelieftochooseactionsandreplan.WefurtherintroduceEvoWorld-Bench,abenchmarkgroundedinhumanactivitytraces,comprising54scenesand803,680taskswithcontrolledchangesbeforeandduringnavigation.Insimulationandreal-robotexperiments,EvolvingNavimprovesnavigationsuccessandsearchefficiencyovertheevaluatedbaselines.Pairedexperimentsshowtheclearestgainsunderlearnabletemporalpatterns,whileablationsdemonstratethevalueofpreservinguncertaintyandincorporatingvisibility-awareevidence.
View arXiv pageView PDFProject pageGitHub0Add to collection
Get this paper in your agent:
hf papers read 2609\.39166
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.39166 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.39166 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.39166 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Generative-Model Predictive Planning for Navigation in Partially Observable Environments
This paper introduces BeliefDiffusion, a framework combining diffusion models to represent multimodal belief distributions and Model Predictive Control for planning in partially observable environments, achieving better navigation success and path efficiency than baselines.
NavHarness: Towards Lifelong Embodied Navigation
The paper introduces NavHarness, a training-free embodied navigation harness that integrates memory processing into the navigation loop, enabling agents to carry experience across sessions for lifelong navigation. On GOAT-Bench, it improves success rate by up to 22.6 points over context-only sessions and achieves state-of-the-art results with GPT-6/Astra.
Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation
TAMP-Nav is a unified framework that enhances embodied navigation by aligning vision-language models with 2D visual prompting, selective reasoning with compressed memory, and policy optimization, achieving state-of-the-art performance with high runtime and sample efficiency.
Predicting Consequences and Reinforcing Navigation Policies with Latent World Models
This paper proposes a Latent World Model (LWM) for robot navigation that predicts action-conditioned latent feature compatibility, enabling policy learning from unlabeled video data and reinforcement learning without additional environment interaction, outperforming existing methods.
Scaffolding Foundation Models into Physical-World Agents Pushes the Frontier of Long-Horizon Navigation
NavMCP is an agentic scaffolding framework that integrates vision-language models with navigation foundation models to enable persistent long-horizon physical-world exploration, achieving state-of-the-art results on benchmarks and real-robot tasks.