EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks
Summary
The paper introduces EmbodiedMemory-Bench, a benchmark for evaluating memory in long-horizon embodied tasks, and presents Embodied-Memorizer, an external memory system with an 8B policy model to enhance agent performance.
View Cached Full Text
Cached at: 09/29/26, 04:13 AM
Paper page - EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks
Source: https://huggingface.co/papers/2609.28236
Abstract
Long-horizonembodiedinteractionrequiresagentstoretainandcontinuallyupdateinformationabouttheenvironmentastheyobserve,act,andencounterchange.Yetcurrentagentsstruggletomaintainsuchmemoryreliably.Ouranalysistracesthislimitationtofourkeydeficiencies:weakfine-grainedvisualmemory,unreliabledynamicworld-statetracking,failingtorecordworldstaterevealedbyinteractionoutcomes,andlimitedgeneralizationfrompriorexperience.However,existingbenchmarksdonotdirectlyassessthesememorycapabilitiesduringlong-horizonembodiedinteraction.Toaddressthisgap,weintroduceEmbodiedMemory-Bench(EMem-Bench),comprising2,554interactiveepisodesacrossfourtaskfamilies.EMem-Benchrequiresagentstobuildandupdatememoryfrominteractionhistory,thenuseittocompletealatertaskbyactingintheenvironment.WefurtherpresentEmbodied-Memorizer(EMem),anexternalmemorysystemthatorganizesembodiedexperienceintospatial,event,andscenememories.WealsotrainEMem-8B,an8Bpolicythatmanagesandusesthesememories.Weevaluateadiverserangeofopen-sourceandproprietaryMLLMsandrepresentativemultimodalmemorysystems.Resultsshowthatcurrentmodelsremainweakandunevenacrossthefourchallenges.Undermatchedbackbones,EMemachievesthebestoverallperformanceamongtheevaluatedmemorysystemsandimprovesbothopen-sourceandproprietarymodels,whileEMem-8Bfurtherimprovesoveritsbackbone.Projectpage:https://zju-omniai.github.io/EmbodiedMemoryBench/
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2609\.28236
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.28236 in a model README.md to link it from this page.
Datasets citing this paper1
#### lzLiang/EmbodiedMemoryBench Viewer• Updated4 days ago • 2.55k • 324 • 2
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.28236 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Personalize-then-Store: Benchmarking and Learning Personalized Memory for Long-horizon Agents
This paper introduces PerMemBench, the first benchmark for evaluating personalized memory systems in LLM-based agents, and proposes a session-level storage gating framework that adapts memory policies to individual user contexts.
AgentMemBench: A Systematic Benchmark for Evaluating Long-Term Memory Management Strategies in Conversational AI Agents
AgentMemBench is a systematic benchmark that evaluates five long-term memory management strategies for conversational AI agents across three datasets, finding that external key-value store retrieval dominates on quality but incurs a larger memory footprint.
WorldLines: Benchmarking and Modeling Long-Horizon Stateful Embodied Agents
WorldLines introduces a benchmark for long-horizon embodied household assistance, featuring memory QA and embodied task planning with partial observability, and proposes ObsMem, a visibility-aware memory framework.
LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues
This paper introduces LongMemEval-V2, a benchmark for evaluating long-term memory systems in web agents, along with two memory methods: AgentRunbook-R and AgentRunbook-C.
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory
The paper presents ABot-AgentOS, a general robotic agent operating system with lifelong multi-modal memory, and introduces EmbodiedWorldBench for evaluating long-horizon embodied tasks. It demonstrates significant improvements in task success and memory benchmarks, suggesting that a dedicated agent OS layer enhances execution and persistent memory.