EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks

Hugging Face Daily Papers Papers

Summary

The paper introduces EmbodiedMemory-Bench, a benchmark for evaluating memory in long-horizon embodied tasks, and presents Embodied-Memorizer, an external memory system with an 8B policy model to enhance agent performance.

Long-horizon embodied interaction requires agents to retain and continually update information about the environment as they observe, act, and encounter change. Yet current agents struggle to maintain such memory reliably. Our analysis traces this limitation to four key deficiencies: weak fine-grained visual memory, unreliable dynamic world-state tracking, failing to record world state revealed by interaction outcomes, and limited generalization from prior experience. However, existing benchmarks do not directly assess these memory capabilities during long-horizon embodied interaction. To address this gap, we introduce EmbodiedMemory-Bench (EMem-Bench), comprising 2,554 interactive episodes across four task families. EMem-Bench requires agents to build and update memory from interaction history, then use it to complete a later task by acting in the environment. We further present Embodied-Memorizer (EMem), an external memory system that organizes embodied experience into spatial, event, and scene memories. We also train EMem-8B, an 8B policy that manages and uses these memories. We evaluate a diverse range of open-source and proprietary MLLMs and representative multimodal memory systems. Results show that current models remain weak and uneven across the four challenges. Under matched backbones, EMem achieves the best overall performance among the evaluated memory systems and improves both open-source and proprietary models, while EMem-8B further improves over its backbone. Project page: https://zju-omniai.github.io/EmbodiedMemoryBench/
Original Article
View Cached Full Text

Cached at: 09/29/26, 04:13 AM

Paper page - EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks

Source: https://huggingface.co/papers/2609.28236

Abstract

Long-horizonembodiedinteractionrequiresagentstoretainandcontinuallyupdateinformationabouttheenvironmentastheyobserve,act,andencounterchange.Yetcurrentagentsstruggletomaintainsuchmemoryreliably.Ouranalysistracesthislimitationtofourkeydeficiencies:weakfine-grainedvisualmemory,unreliabledynamicworld-statetracking,failingtorecordworldstaterevealedbyinteractionoutcomes,andlimitedgeneralizationfrompriorexperience.However,existingbenchmarksdonotdirectlyassessthesememorycapabilitiesduringlong-horizonembodiedinteraction.Toaddressthisgap,weintroduceEmbodiedMemory-Bench(EMem-Bench),comprising2,554interactiveepisodesacrossfourtaskfamilies.EMem-Benchrequiresagentstobuildandupdatememoryfrominteractionhistory,thenuseittocompletealatertaskbyactingintheenvironment.WefurtherpresentEmbodied-Memorizer(EMem),anexternalmemorysystemthatorganizesembodiedexperienceintospatial,event,andscenememories.WealsotrainEMem-8B,an8Bpolicythatmanagesandusesthesememories.Weevaluateadiverserangeofopen-sourceandproprietaryMLLMsandrepresentativemultimodalmemorysystems.Resultsshowthatcurrentmodelsremainweakandunevenacrossthefourchallenges.Undermatchedbackbones,EMemachievesthebestoverallperformanceamongtheevaluatedmemorysystemsandimprovesbothopen-sourceandproprietarymodels,whileEMem-8Bfurtherimprovesoveritsbackbone.Projectpage:https://zju-omniai.github.io/EmbodiedMemoryBench/

View arXiv pageView PDFProject pageAdd to collection

Get this paper in your agent:

hf papers read 2609\.28236

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.28236 in a model README.md to link it from this page.

Datasets citing this paper1

#### lzLiang/EmbodiedMemoryBench Viewer• Updated4 days ago • 2.55k • 324 • 2

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.28236 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory

Hugging Face Daily Papers

The paper presents ABot-AgentOS, a general robotic agent operating system with lifelong multi-modal memory, and introduces EmbodiedWorldBench for evaluating long-horizon embodied tasks. It demonstrates significant improvements in task success and memory benchmarks, suggesting that a dedicated agent OS layer enhances execution and persistent memory.