Spatial Memory Intelligence: Endowing World Models with Understanding-Driven Long-Term Memory

Hugging Face Daily Papers Papers

Summary

The paper proposes Spatial Memory Intelligence (SMI), the first framework to use an understanding model (MLLM) to systematically manage spatial memory in long-video world models via four atomic operations: spatial clustering, within-cluster sparsification, action-aware retrieval, and reliability-aware filtering. Experiments show improvements in memory sparsity, generation stability, and spatial consistency across multiple world-model backbones.

Long-video generation and world models have shown strong potential for interactive entertainment and embodied simulation by predicting future observations conditioned on user actions and historical memory. However, as memory sequences grow longer and their structures become increasingly complex, managing long-range spatial context becomes increasingly challenging, calling for a more intelligent and systematic memory-management strategy. Building on the advancing spatial reasoning capabilities of multimodal large language models (MLLMs) and the broader vision of unified models, we propose Spatial Memory Intelligence (SMI), the first framework to systematically employ an understanding model for spatial-memory management in long-video world models. SMI introduces four coordinated atomic operations: spatial clustering, within-cluster sparsification, action-aware retrieval, and reliability-aware filtering. Extensive experiments across multiple baselines, benchmarks, and world-model backbones demonstrate the effectiveness and generalizability of SMI, achieving comprehensive improvements in memory sparsity, generation stability, and spatial consistency.
Original Article
View Cached Full Text

Cached at: 10/05/26, 08:46 AM

Paper page - Spatial Memory Intelligence: Endowing World Models with Understanding-Driven Long-Term Memory

Source: https://huggingface.co/papers/2610.02521

Abstract

Long-videogenerationandworldmodelshaveshownstrongpotentialforinteractiveentertainmentandembodiedsimulationbypredictingfutureobservationsconditionedonuseractionsandhistoricalmemory.However,asmemorysequencesgrowlongerandtheirstructuresbecomeincreasinglycomplex,managinglong-rangespatialcontextbecomesincreasinglychallenging,callingforamoreintelligentandsystematicmemory-managementstrategy.Buildingontheadvancingspatialreasoningcapabilitiesofmultimodallargelanguagemodels(MLLMs)andthebroadervisionofunifiedmodels,weproposeSpatialMemoryIntelligence(SMI),thefirstframeworktosystematicallyemployanunderstandingmodelforspatial-memorymanagementinlong-videoworldmodels.SMIintroducesfourcoordinatedatomicoperations:spatialclustering,within-clustersparsification,action-awareretrieval,andreliability-awarefiltering.Extensiveexperimentsacrossmultiplebaselines,benchmarks,andworld-modelbackbonesdemonstratetheeffectivenessandgeneralizabilityofSMI,achievingcomprehensiveimprovementsinmemorysparsity,generationstability,andspatialconsistency.

View arXiv pageView PDFProject pageGitHubAdd to collection

Get this paper in your agent:

hf papers read 2610\.02521

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2610.02521 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2610.02521 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2610.02521 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Latent Spatial Memory for Video World Models

Hugging Face Daily Papers

This paper introduces latent spatial memory for video world models, storing 3D scene information directly in diffusion latent space to avoid costly pixel-space reconstruction. The proposed Mirage framework achieves up to 10.57x faster generation and 55x memory reduction while achieving state-of-the-art performance on WorldScore and RealEstate10K.

Mental World Modeling

Hugging Face Daily Papers

The paper introduces Mental World Modeling (MWM), a framework that integrates hidden mental states as core components of world models, and presents MENTIS, a training-free baseline. Experiments with 8 LLM-based world models show explicit mental-state modeling is essential for predicting human decisions in situated scenarios.

Composition of Memory Experts for Diffusion World Models

arXiv cs.LG

A new diffusion-based world model framework that uses a composition of specialized memory experts (short-term, long-term episodic, and spatial) to achieve better temporal consistency and long-context modeling without quadratic cost.

When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents

Hugging Face Daily Papers

This paper empirically studies spatial memory staleness in vision-language-model agents, finding that models often ignore contradictory visual evidence and that trusting stale memory can increase safety risks. The authors propose auditing mechanisms but show that visual grounding under memory-observation conflicts remains a major open challenge.