EpiCon: Collective Agent Learning through Co-Evolving Multimodal Memory

Hugging Face Daily Papers Papers

Summary

EpiCon presents a shared multimodal memory framework for collective learning among AI agents, enhancing performance across eleven benchmarks without updating host model parameters.

Agents can learn from past executions, but enabling different agents to reuse and build on one another's experience remains challenging. We introduce EpiCon, a shared multimodal memory framework for agent collective learning without updating host model parameters. EpiCon links question-level memory evolution to a persistent experience bank through two independently trained 2B models: a memory controller and a tree self-organizer. The controller jointly refines textual guidance and visual evidence across attempts and selectively includes visual memory. The self-organizer consolidates lessons hierarchically and retrieves experience and rules for new problems. We evaluate EpiCon on eleven benchmarks spanning four multimodal task domains, using two harnesses and multiple backbones. A frozen bank improves other systems even with a single solving attempt. A second harness raises the original system's macro-average score by 2.6 points across eleven benchmarks. Across four host configurations, EpiCon improves macro-average scores by 1.7 to 4.9 points over No Memory and reduces memory-operation time by 67\% to 74\% relative to backbone-sized memory models.
Original Article
View Cached Full Text

Cached at: 09/30/26, 08:23 AM

Paper page - EpiCon: Collective Agent Learning through Co-Evolving Multimodal Memory

Source: https://huggingface.co/papers/2609.37923

Abstract

Agentscanlearnfrompastexecutions,butenablingdifferentagentstoreuseandbuildononeanother’sexperienceremainschallenging.WeintroduceEpiCon,asharedmultimodalmemoryframeworkforagentcollectivelearningwithoutupdatinghostmodelparameters.EpiConlinksquestion-levelmemoryevolutiontoapersistentexperiencebankthroughtwoindependentlytrained2Bmodels:amemorycontrollerandatreeself-organizer.Thecontrollerjointlyrefinestextualguidanceandvisualevidenceacrossattemptsandselectivelyincludesvisualmemory.Theself-organizerconsolidateslessonshierarchicallyandretrievesexperienceandrulesfornewproblems.WeevaluateEpiCononelevenbenchmarksspanningfourmultimodaltaskdomains,usingtwoharnessesandmultiplebackbones.Afrozenbankimprovesothersystemsevenwithasinglesolvingattempt.Asecondharnessraisestheoriginalsystem’smacro-averagescoreby2.6pointsacrosselevenbenchmarks.Acrossfourhostconfigurations,EpiConimprovesmacro-averagescoresby1.7to4.9pointsoverNoMemoryandreducesmemory-operationtimeby67\%to74\%relativetobackbone-sizedmemorymodels.

View arXiv pageView PDFProject pageAdd to collection

Get this paper in your agent:

hf papers read 2609\.37923

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper1

#### zengziyun/EpiCon Updatedabout 4 hours ago

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.37923 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.37923 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments

Hugging Face Daily Papers

EvoArena introduces a benchmark for evaluating LLM agents in dynamic environments with progressive updates across terminal, software, and social domains, while EvoMem proposes a patch-based memory paradigm that records structured evolution; experiments show current agents achieve only 39.6% accuracy on EvoArena, and EvoMem yields average gains of 1.5% on the benchmark and improvements on GAIA and LoCoMo.

Collaborative Memory for Multi-Agent VLM Systems

arXiv cs.AI

This paper proposes a collaborative memory framework for multi-agent vision-language model systems to address distributed perception and improve shared visual context and reasoning consistency.