Tag
This paper shows that for coding agents' memory systems, checking whether a specific claim remains valid after a repository change provides higher precision than assessing behavior preservation in diffs, validated through experiments with multiple LLMs and real-world data.
This paper empirically studies how VLM agents with persistent spatial memory fail when memory becomes stale, using a dynamic FrozenLake testbed. It finds that trusting stale memory can more than double death rates, and that read-time auditing helps but does not fully close the gap.
This paper empirically studies spatial memory staleness in vision-language-model agents, finding that models often ignore contradictory visual evidence and that trusting stale memory can increase safety risks. The authors propose auditing mechanisms but show that visual grounding under memory-observation conflicts remains a major open challenge.