Spatial Memory Intelligence: Endowing World Models with Understanding-Driven Long-Term Memory
Summary
The paper proposes Spatial Memory Intelligence (SMI), the first framework to use an understanding model (MLLM) to systematically manage spatial memory in long-video world models via four atomic operations: spatial clustering, within-cluster sparsification, action-aware retrieval, and reliability-aware filtering. Experiments show improvements in memory sparsity, generation stability, and spatial consistency across multiple world-model backbones.
View Cached Full Text
Cached at: 10/05/26, 08:46 AM
Paper page - Spatial Memory Intelligence: Endowing World Models with Understanding-Driven Long-Term Memory
Source: https://huggingface.co/papers/2610.02521
Abstract
Long-videogenerationandworldmodelshaveshownstrongpotentialforinteractiveentertainmentandembodiedsimulationbypredictingfutureobservationsconditionedonuseractionsandhistoricalmemory.However,asmemorysequencesgrowlongerandtheirstructuresbecomeincreasinglycomplex,managinglong-rangespatialcontextbecomesincreasinglychallenging,callingforamoreintelligentandsystematicmemory-managementstrategy.Buildingontheadvancingspatialreasoningcapabilitiesofmultimodallargelanguagemodels(MLLMs)andthebroadervisionofunifiedmodels,weproposeSpatialMemoryIntelligence(SMI),thefirstframeworktosystematicallyemployanunderstandingmodelforspatial-memorymanagementinlong-videoworldmodels.SMIintroducesfourcoordinatedatomicoperations:spatialclustering,within-clustersparsification,action-awareretrieval,andreliability-awarefiltering.Extensiveexperimentsacrossmultiplebaselines,benchmarks,andworld-modelbackbonesdemonstratetheeffectivenessandgeneralizabilityofSMI,achievingcomprehensiveimprovementsinmemorysparsity,generationstability,andspatialconsistency.
View arXiv pageView PDFProject pageGitHubAdd to collection
Get this paper in your agent:
hf papers read 2610\.02521
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2610.02521 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2610.02521 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2610.02521 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence
This paper introduces Spatial Memory Agent (SMA), a runtime framework that improves frozen vision-language models' spatial reasoning through verifier-guided reflection and reusable memory without parameter updates or external tools, achieving strong results across five benchmarks and four base VLMs.
Latent Spatial Memory for Video World Models
This paper introduces latent spatial memory for video world models, storing 3D scene information directly in diffusion latent space to avoid costly pixel-space reconstruction. The proposed Mirage framework achieves up to 10.57x faster generation and 55x memory reduction while achieving state-of-the-art performance on WorldScore and RealEstate10K.
Mental World Modeling
The paper introduces Mental World Modeling (MWM), a framework that integrates hidden mental states as core components of world models, and presents MENTIS, a training-free baseline. Experiments with 8 LLM-based world models show explicit mental-state modeling is essential for predicting human decisions in situated scenarios.
Composition of Memory Experts for Diffusion World Models
A new diffusion-based world model framework that uses a composition of specialized memory experts (short-term, long-term episodic, and spatial) to achieve better temporal consistency and long-context modeling without quadratic cost.
When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents
This paper empirically studies spatial memory staleness in vision-language-model agents, finding that models often ignore contradictory visual evidence and that trusting stale memory can increase safety risks. The authors propose auditing mechanisms but show that visual grounding under memory-observation conflicts remains a major open challenge.