When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents
Summary
This paper empirically studies spatial memory staleness in vision-language-model agents, finding that models often ignore contradictory visual evidence and that trusting stale memory can increase safety risks. The authors propose auditing mechanisms but show that visual grounding under memory-observation conflicts remains a major open challenge.
View Cached Full Text
Cached at: 08/06/26, 05:50 AM
Paper page - When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents
Source: https://huggingface.co/papers/2608.04574
Abstract
Memory-augmentedVLMagentsactonpersistentspatialknowledge,yetthatknowledgesilentlygoesstaleastheenvironmentchanges.Weaskwhathappenswhenanagentmustreconcileaconfidentmemoryclaimwithacontradictingobservation,andwhethercurrentmodelscancatchtheconflictbeforeitbecomesasafety-relevantmistake.UsingadynamicFrozenLaketestbed,wepairastaleness-detectiontaskwithadownstreamnavigationtaskacrossthreeclosed-sourcemodelsandthreeopen-weightVLMsunderbothtextandimageinputs(1,800detectionruns,and12,000text-modenavigationepisodesoverfourLLMnavigatorsatashared50-seedscale).Threefindingsemerge.First,textsolvabilitydoesnotimplyvisualgrounding:modelsthatflagstaleentriesreliablyfromtextnonethelessspanvisionF1from0.887downto0.067ontheidenticalgrids,andtheweakestkeepsmakingfluent,confidentdecisionsthatignoretheimage.Second,consumingstalememorywithoutanauditisasafetyliability:inourprimaryGPT-4osetting,anagentthattrustsrawmemorydiesmorethantwiceasoftenasthesameagentgivennomemoryatall.Third,auditinghelpsbutdoesnotclosethegap:atransparentread-timefilterremovesmuchofthesafetycostintextmode,yetevenoraclestalelabelsbringnofurthersignificantgainonthecurrentgridsize,andwhenvisualauditingisunreliable,filteringyieldsnoconsistentbenefit.Togethertheseresultsframespatial-memorystalenessasasafetyfailuremodeandisolatereliablevisualgroundingandactionselectionundermemory--observationconflictasthecentralopenchallengesformemory-augmentedagents.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2608\.04574
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.04574 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2608.04574 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.04574 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents
This paper empirically studies how VLM agents with persistent spatial memory fail when memory becomes stale, using a dynamic FrozenLake testbed. It finds that trusting stale memory can more than double death rates, and that read-time auditing helps but does not fully close the gap.
STALE: Can LLM Agents Know When Their Memories Are No Longer Valid?
This paper identifies a critical failure mode in LLM agents where they fail to update personalized memories when new evidence conflicts with prior beliefs. It introduces the STALE benchmark and a three-dimensional probing framework, revealing that even the best models achieve only 55.2% accuracy, and proposes CUPMem as a prototype for robust memory revision.
What Spatial Memory Must Store: Occlusion as the Test for Language-Agent Memory
This paper investigates whether spatial geometry improves language-agent memory recall, demonstrating that geometry must lead recall over recency and importance, and that a ray-tracing visibility predicate is crucial for occlusion handling in 3D voxel worlds.
Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence
This paper introduces Spatial Memory Agent (SMA), a runtime framework that improves frozen vision-language models' spatial reasoning through verifier-guided reflection and reusable memory without parameter updates or external tools, achieving strong results across five benchmarks and four base VLMs.
Collaborative Memory for Multi-Agent VLM Systems
This paper proposes a collaborative memory framework for multi-agent vision-language model systems to address distributed perception and improve shared visual context and reasoning consistency.