EvoGraph-Mem: Failure-Aware Editable Graph Memory for Long-Term Language Agents

arXiv cs.AI Papers

Summary

Presents EvoGraph-Mem, a failure-aware editable graph memory framework for long-term language agents that tracks positive/negative evidence and activation states for insights, enabling memory maintenance through utility-aware retrieval and graph-level editing.

arXiv:2608.11248v1 Announce Type: new Abstract: Long-term memory is essential for language agents operating across extended interactions and evolving tasks. Existing memory-augmented agents mainly focus on storing and retrieving past experience, but the quality of stored memories may degrade over time. In particular, previously distilled insights can become outdated, over-generalized, or harmful under new task contexts, causing memory pollution when repeatedly reused. To address this issue, we study insight-level memory maintenance for long-term language agents and propose a failure-aware memory maintenance framework based on an editable insight graph. Each insight node tracks positive evidence, negative evidence, and an activation state, enabling the agent to distinguish reusable insights from conflicting or invalid ones. We further introduce a utility-aware retrieval mechanism and a graph controller that updates the memory graph after task execution by keeping reliable insights, archiving invalid ones, revising outdated ones, and adding newly discovered reusable insights. Extensive experiments show that our method consistently outperforms representative memory-based agent baselines across different backbone models. Ablation studies further demonstrate that append-only memory is insufficient for long-horizon tasks, while evidence-aware retrieval and graph-level editing improve memory reliability and downstream task performance.
Original Article
View Cached Full Text

Cached at: 08/13/26, 03:24 PM

# EvoGraph-Mem: Failure-Aware Editable Graph Memory for Long-Term Language Agents
Source: [https://arxiv.org/html/2608.11248](https://arxiv.org/html/2608.11248)
###### Abstract

Long\-term memory is essential for language agents operating across extended interactions and evolving tasks\. Existing memory\-augmented agents mainly focus on storing and retrieving past experience, but the quality of stored memories may degrade over time\. In particular, previously distilled insights can become outdated, over\-generalized, or harmful under new task contexts, causing memory pollution when repeatedly reused\. To address this issue, we study insight\-level memory maintenance for long\-term language agents and propose a failure\-aware memory maintenance framework based on an editable insight graph\. Each insight node tracks positive evidence, negative evidence, and an activation state, enabling the agent to distinguish reusable insights from conflicting or invalid ones\. We further introduce a utility\-aware retrieval mechanism and a graph controller that updates the memory graph after task execution by keeping reliable insights, archiving invalid ones, revising outdated ones, and adding newly discovered reusable insights\. Extensive experiments show that our method consistently outperforms representative memory\-based agent baselines across different backbone models\. Ablation studies further demonstrate that append\-only memory is insufficient for long\-horizon tasks, while evidence\-aware retrieval and graph\-level editing improve memory reliability and downstream task performance\.

EvoGraph\-Mem: Failure\-Aware Editable Graph Memory for Long\-Term Language Agents

Yuxi QianYuxiang Ren\*

## 1Introduction

Large language model \(LLM\) agents are increasingly expected to operate across extended interactions, multi\-step tasks, and evolving environments\. Long\-term memory is therefore essential for retaining useful experience and reusing it in future decision making\. Existing memory\-augmented agents store past interactions, reflections, or reusable skills to improve long\-horizon behavior\(Zhonget al\.,[2024](https://arxiv.org/html/2608.11248#bib.bib7); Wanget al\.,[2023](https://arxiv.org/html/2608.11248#bib.bib8)\)\. However, long\-term memory is not only a problem of storage and retrieval, but also one of maintaining memory quality over time\.

As agents encounter diverse tasks, previously distilled insights may become outdated, over\-generalized, or harmful under new contexts\. Once such insights are repeatedly retrieved, they may introduce memory pollution and degrade downstream reasoning\. This issue is especially important for high\-level insights, which are intended to generalize across tasks but may cause greater harm when applied beyond their valid contexts\. Recent graph\-based memory systems, such as G\-Memory, improve memory organization by structuring historical experience into insight, query, and interaction graphs\(Zhanget al\.,[2025](https://arxiv.org/html/2608.11248#bib.bib1)\)\. Nevertheless, existing memory systems remain limited in explicitly modeling whether a recalled insight was helpful, conflicting, or should be revised after task failure\.

In this paper, we study*insight\-level memory maintenance*for long\-term language agents and propose a failure\-aware memory maintenance framework based on an editable insight graph\. Each insight node explicitly tracks positive evidence, negative evidence, and an activation state: positive evidence records queries for which the insight provides useful guidance, negative evidence captures cases where the insight is ineffective or harmful, and the activation state distinguishes active insights from archived ones to prevent invalid memories from being repeatedly retrieved\. Building on this representation, we introduce utility\-aware retrieval and a graph controller for memory correction\. Candidate insights are ranked by jointly considering positive support and conflicting evidence, while after task execution, the graph controller updates the insight graph by keeping reliable insights, archiving invalid ones, revising outdated ones, and adding new insights when reusable knowledge is discovered\. Unlike model editing methods that modify internal parameters\(Mitchellet al\.,[2021](https://arxiv.org/html/2608.11248#bib.bib25)\), our method operates at the external agent\-memory level and keeps the underlying LLM unchanged\.

Extensive experiments show that our framework consistently improves over representative memory\-based agent methods\. Ablation studies further demonstrate that append\-only memory is insufficient, while evidence\-aware retrieval and graph\-level editing contribute to stronger long\-horizon performance\.

Our contributions are summarized as follows:

- •Identify insight\-level memory maintenance as a challenge for long\-term language agents, where useful insights may later become over\-generalized, conflicting, or harmful\.
- •Propose an editable insight graph that tracks positive evidence, negative evidence, and activation states for failure\-aware retrieval\.
- •Introduce a graph controller that performs corrective memory updates through keeping, archiving, revising, and adding insights\.
- •Experiments demonstrate consistent improvements over representative memory\-based agent baselines, with ablations validating the importance of graph\-level editing\.

![Refer to caption](https://arxiv.org/html/2608.11248v1/x1.png)Figure 1:The mechanism of memory graph update
## 2Related Work

### 2\.1Long\-Term Memory for LLM Agents

Long\-term memory has become a key component for language agents that operate across extended interactions, tasks, or environments\. Existing systems represent persistent experience in different forms, including conversational histories, episodic memories, reflective summaries, and executable skills\. Generative Agents maintain memory streams and synthesize higher\-level reflections to support coherent interactive behavior\(Parket al\.,[2023](https://arxiv.org/html/2608.11248#bib.bib9)\)\. Reflexion converts task feedback into verbal reflections stored in episodic memory, enabling agents to improve later decisions without parameter updates\(Shinnet al\.,[2023](https://arxiv.org/html/2608.11248#bib.bib10)\)\. MemoryBank studies long\-term conversational memory for recalling and updating past interactions\(Zhonget al\.,[2024](https://arxiv.org/html/2608.11248#bib.bib7)\), while Voyager accumulates reusable skills for open\-ended embodied exploration\(Wanget al\.,[2023](https://arxiv.org/html/2608.11248#bib.bib8)\)\. MemGPT further frames memory management as virtual context management across memory tiers\(Packeret al\.,[2023](https://arxiv.org/html/2608.11248#bib.bib11)\)\. Recent systems such as MemoryOS, Mem0, and A\-MEM move beyond simple storage by emphasizing hierarchical organization, dynamic updating, scalable retrieval, and adaptive memory networks\(Kanget al\.,[2025](https://arxiv.org/html/2608.11248#bib.bib18); Chhikaraet al\.,[2025](https://arxiv.org/html/2608.11248#bib.bib17); Xuet al\.,[2026](https://arxiv.org/html/2608.11248#bib.bib16)\)\. These studies show that persistent memory is essential for long\-horizon agents\. However, they largely focus on how memories are stored, retrieved, summarized, or expanded, while the validity of a distilled insight after repeated reuse remains less explicitly modeled\. Our work complements this direction by focusing on insight\-level memory maintenance, where retrieved memories are evaluated after task execution and may be reinforced, suppressed, or revised\.

### 2\.2Graph Memory and Corrective Knowledge Maintenance

Graph structures provide a natural mechanism for organizing relational dependencies among tasks, experiences, and abstract knowledge\. In retrieval\-augmented generation, HippoRAG combines language models, knowledge graphs, and graph\-based retrieval to support long\-term knowledge integration and multi\-hop reasoning\(Gutiérrezet al\.,[2024](https://arxiv.org/html/2608.11248#bib.bib24)\)\. For agent memory, AriGraph integrates semantic and episodic memories into a graph\-based world model\(Anokhinet al\.,[2024](https://arxiv.org/html/2608.11248#bib.bib20)\), and Zep introduces a temporal knowledge graph architecture for continuously evolving conversational memory\(Rasmussenet al\.,[2025](https://arxiv.org/html/2608.11248#bib.bib21)\)\. Closely related to our work, G\-Memory organizes multi\-agent histories into a hierarchical graph consisting of insight, query, and interaction graphs, enabling agents to retrieve both high\-level insights and fine\-grained interaction trajectories\(Zhanget al\.,[2025](https://arxiv.org/html/2608.11248#bib.bib1)\)\. These graph\-based approaches improve memory organization and retrieval, but they remain limited in explicitly handling cases where a previously useful insight becomes over\-generalized, conflicting, or harmful under new task contexts\.

Our work is closely related to knowledge editing, which seeks to correct outdated or erroneous knowledge in language models without full retraining\. Existing approaches, such as MEND, ROME, MEMIT, and GRACE, modify model parameters or auxiliary memory structures to update factual associations while preserving the model’s general behavior\(Mitchellet al\.,[2021](https://arxiv.org/html/2608.11248#bib.bib25); Menget al\.,[2022a](https://arxiv.org/html/2608.11248#bib.bib26),[b](https://arxiv.org/html/2608.11248#bib.bib27); Hartvigsenet al\.,[2023](https://arxiv.org/html/2608.11248#bib.bib28)\)\. Subsequent methods, including AlphaEdit and AdaEdit, further examine how edited knowledge can be preserved under sequential or continuous editing scenarios\(Fanget al\.,[2025](https://arxiv.org/html/2608.11248#bib.bib22); Li and Chu,[2025](https://arxiv.org/html/2608.11248#bib.bib23)\)\. Despite this shared objective, our work differs in that it addresses unreliable knowledge at the external agent\-memory level, rather than through parameter\-level intervention\. We introduce an editable insight graph in which nodes record positive evidence, negative evidence, and activation states\. This structure enables the system to archive invalid insights, update outdated ones, and mitigate memory pollution during long\-term agent deployment\.

## 3Methodology

This section presents our failure\-aware memory maintenance framework for long\-term agent memory, as illustrated in Figure[1](https://arxiv.org/html/2608.11248#S1.F1)\. Given a new user queryQQ, the agent first retrieves relevant historical queries and insights following the graph\-based memory structure of G\-memory\(Zhanget al\.,[2025](https://arxiv.org/html/2608.11248#bib.bib1)\)\. However, unlike append\-only memory systems, our framework explicitly models whether a recalled insight remains useful, becomes harmful, or requires revision after being applied to a new task\.

Specifically, we extend each insight node with positive evidence, negative evidence, and an activation state, enabling the system to distinguish reusable insights from non\-generalizable or polluted ones\. After task completion, a graph controller updates the insight graph through three maintenance operations: keeping reliable insights, archiving invalid insights, and revising outdated insights\. This design allows the memory graph to evolve from raw accumulation toward corrective maintenance, thereby reducing the long\-term propagation of invalid or harmful memories\.

### 3\.1Graph\-Based Historical Memory

During the initial forward pass, the query graph is defined as:

𝒢query=\(𝒬,ℰq\)=\(\{Qi,Ψi,Ginter\(Qi\)\}i=1\|𝒬\|,ℰq\)\.\\mathcal\{G\}\_\{\\text\{query\}\}=\(\\mathcal\{Q\},\\mathcal\{E\}\_\{q\}\)=\\left\(\\left\\\{Q\_\{i\},\\Psi\_\{i\},G^\{\(Q\_\{i\}\)\}\_\{\\text\{inter\}\}\\right\\\}\_\{i=1\}^\{\|\\mathcal\{Q\}\|\},\\mathcal\{E\}\_\{q\}\\right\)\.\(1\)Here,𝒬=\{qi\}\\mathcal\{Q\}=\\\{q\_\{i\}\\\}denotes the set of query nodes, where each nodeqi≜\(Qi,Ψi,Ginter\(Qi\)\)q\_\{i\}\\triangleq\(Q\_\{i\},\\Psi\_\{i\},G^\{\(Q\_\{i\}\)\}\_\{\\text\{inter\}\}\)consists of the original queryQiQ\_\{i\}, its task statusΨi∈\{Failed,Resolved\}\\Psi\_\{i\}\\in\\\{\\text\{Failed\},\\text\{Resolved\}\\\}, and the corresponding interaction graphGinter\(Qi\)G^\{\(Q\_\{i\}\)\}\_\{\\text\{inter\}\}\. The edge setℰq⊆𝒬×𝒬\\mathcal\{E\}\_\{q\}\\subseteq\\mathcal\{Q\}\\times\\mathcal\{Q\}represents semantic dependencies between queries\. Leveraging this structured topology, the query graph supports more precise retrieval compared to coarse similarity\-based approaches such as embedding matching\.

Upon receiving a new queryQQ, the multi\-agent system performs a retrieval operation on the graph𝒢query\\mathcal\{G\}\_\{\\text\{query\}\}maintained from historical tasks to identify and recall analogous instances:

𝒬𝒮=arg⁡top\-kqi∈𝒬​s\.t\.​\|𝒬𝒮\|=k​\(𝐯​\(Q\)⋅𝐯​\(qi\)\|𝐯​\(Q\)\|​\|𝐯​\(qi\)\|\)\.\\mathcal\{Q\}^\{\\mathcal\{S\}\}=\\underset\{q\_\{i\}\\in\\mathcal\{Q\}\\text\{ s\.t\. \}\\left\|\\mathcal\{Q\}^\{\\mathcal\{S\}\}\\right\|=k\}\{\\arg\\text\{ top\-k \}\}\\left\(\\frac\{\\mathbf\{v\}\(Q\)\\cdot\\mathbf\{v\}\\left\(q\_\{i\}\\right\)\}\{\|\\mathbf\{v\}\(Q\)\|\\left\|\\mathbf\{v\}\\left\(q\_\{i\}\\right\)\\right\|\}\\right\)\.\(2\)In this formulation,𝐯​\(⋅\)\\mathbf\{v\}\(\\cdot\)denotes the vectorization applied to both the historical queries and the current query\. By computing the cosine similarity between these representations, the system can retrieve historical queries that exhibit high semantic similarity to the current one\. However, relying exclusively on this metric may introduce noise\. To mitigate this, corresponding extension techniques are further incorporated:

Δ​𝒬=\{Qk∈𝒬∣∃Qj∈𝒬𝒮​s\.t\.​Qk∈𝒩​\(Qj\)\}\.\\Delta\\mathcal\{Q\}=\\\{\{Q\}\_\{k\}\\in\\mathcal\{Q\}\\mid\\exists Q\_\{j\}\\in\\mathcal\{Q\}^\{\\mathcal\{S\}\}\\text\{ s\.t\. \}Q\_\{k\}\\in\\mathcal\{N\}\(Q\_\{j\}\)\\\}\.\(3\)𝒬~𝒮=𝒬𝒮∪Δ​𝒬\.\\tilde\{\\mathcal\{Q\}\}^\{\\mathcal\{S\}\}=\\mathcal\{Q\}^\{\\mathcal\{S\}\}\\cup\\Delta\\mathcal\{Q\}\.\(4\)Here,Δ​𝒬\\Delta\\mathcal\{Q\}denotes the set of neighboring nodes for all nodes within𝒬𝒮\\mathcal\{Q\}^\{\\mathcal\{S\}\}\. Through the expansion of the original𝒬~𝒮\\tilde\{\\mathcal\{Q\}\}^\{\\mathcal\{S\}\}, a task\-specific subgraph is subsequently constructed based on these retrieved elements\.

The insight graph𝒢insight\\mathcal\{G\}\_\{\\text\{insight\}\}is defined as:

𝒢insight=\(ℐ,ℰi\)=\(⟨κk,Ωk⟩k=1\|ℐ\|,ℰi\)\.\\mathcal\{G\}\_\{\\text\{insight\}\}=\(\\mathcal\{I\},\\mathcal\{E\}\_\{\\text\{i\}\}\)=\\left\(\\left\\langle\\kappa\_\{k\},\\Omega\_\{k\}\\right\\rangle\_\{k=1\}^\{\|\\mathcal\{I\}\|\},\\mathcal\{E\}\_\{\\text\{i\}\}\\right\)\.\(5\)Here, the node setℐ=\{ιk\}\\mathcal\{I\}=\\\{\\iota\_\{k\}\\\}denotes the distilled insights, with each nodeιk\\iota\_\{k\}comprising the insight contentκk\\kappa\_\{k\}and an associated set of supporting queriesΩk⊆𝒬\\Omega\_\{k\}\\subseteq\\mathcal\{Q\}\. Furthermore, the edge setℰi⊆ℐ×ℐ×𝒬\\mathcal\{E\}\_\{\\text\{i\}\}\\subseteq\\mathcal\{I\}\\times\\mathcal\{I\}\\times\\mathcal\{Q\}establishes hyper\-connections, wherein a tuple\(ιm,ιn,qj\)\(\\iota\_\{m\},\\iota\_\{n\},q\_\{j\}\)indicates that insightιm\\iota\_\{m\}contextualizesιn\\iota\_\{n\}, mediated by queryqjq\_\{j\}\.

### 3\.2Failure\-Aware Insight Representation

A central limitation of existing insight memories is that they typically store only the queries that contributed to the creation of an insight\. Such a representation implicitly assumes that an insight remains valid once generated\. In long\-term agent deployment, however, this assumption can lead to memory pollution: an insight that was useful for one query may be repeatedly retrieved for semantically similar but incompatible tasks, thereby degrading future reasoning\.

To make insight validity explicitly controllable, we represent each insight node as

ιk=\(κk,Ωk\+,Ωk−,zk\),\\iota\_\{k\}=\(\\kappa\_\{k\},\\Omega\_\{k\}^\{\+\},\\Omega\_\{k\}^\{\-\},z\_\{k\}\),\(6\)whereκk\\kappa\_\{k\}denotes the insight content,Ωk\+\\Omega\_\{k\}^\{\+\}denotes the set of queries for which the insight provides positive support,Ωk−\\Omega\_\{k\}^\{\-\}denotes the set of queries for which the insight is ineffective or harmful, andzk∈\{0,1\}z\_\{k\}\\in\\\{0,1\\\}denotes the activation state of the node\. An active node \(zk=1z\_\{k\}=1\) can be retrieved for future memory augmentation, whereas an archived node \(zk=0z\_\{k\}=0\) is excluded from retrieval but retained for auditability and edit tracing\.

### 3\.3Utility\-Aware Insight Retrieval

Given the expanded query subgraph𝒬~𝒮\\tilde\{\\mathcal\{Q\}\}^\{\\mathcal\{S\}\}, we retrieve insights that are both active and supported by at least one structurally related historical query:

Itcand=\{ιk∣zk=1,Ωk\+∩𝒬~𝒮≠∅\}\.I\_\{t\}^\{\\mathrm\{cand\}\}=\\left\\\{\\iota\_\{k\}\\mid z\_\{k\}=1,\\;\\Omega\_\{k\}^\{\+\}\\cap\\tilde\{\\mathcal\{Q\}\}^\{\\mathcal\{S\}\}\\neq\\varnothing\\right\\\}\.\(7\)
However, positive relevance alone is insufficient: an insight may be useful for some historical queries while conflicting with others\. To suppress polluted or non\-generalizable insights, we score each candidate using both positive and negative evidence:

rt,k=st,k−λ​ct,k,r\_\{t,k\}=s\_\{t,k\}\-\\lambda c\_\{t,k\},\(8\)where

st,k=\|Ωk\+∩𝒬~𝒮\|,ct,k=\|Ωk−∩𝒬~𝒮\|\.s\_\{t,k\}=\|\\Omega\_\{k\}^\{\+\}\\cap\\tilde\{\\mathcal\{Q\}\}^\{\\mathcal\{S\}\}\|,\\quad c\_\{t,k\}=\|\\Omega\_\{k\}^\{\-\}\\cap\\tilde\{\\mathcal\{Q\}\}^\{\\mathcal\{S\}\}\|\.\(9\)Here,st,ks\_\{t,k\}measures the amount of relevant positive evidence, whilect,kc\_\{t,k\}measures the amount of conflicting historical evidence\. The coefficientλ≥0\\lambda\\geq 0controls the strength of conflict penalization\. We then select the top\-BBinsights:

Itret=TopBιk∈Itcand⁡rt,k\.I\_\{t\}^\{\\mathrm\{ret\}\}=\\operatorname\{TopB\}\_\{\\iota\_\{k\}\\in I\_\{t\}^\{\\mathrm\{cand\}\}\}r\_\{t,k\}\.\(10\)

### 3\.4Graph Controller for Memory Correction

After the agent completes the current task, the retrieved insights are no longer treated as static memories\. Instead, the graph controller evaluates whether each insight retrieved was helpful, harmful, or insufficient for the current query\. This post\-task feedback enables corrective memory maintenance\. We instantiate the graph controller as a prompt\-constrained LLM\. Given the current query, execution trajectory, task feedback, and retrieved insight nodes, the controller is required to output a valid object whose operations are restricted to KEEP, ARCHIVE, and REVISE for existing insights, with a separate binary ADD decision for new insight creation\.

For each retrieved insightιk∈Itret\\iota\_\{k\}\\in I\_\{t\}^\{\\mathrm\{ret\}\}, the controller predicts an edit operation:

ok\\displaystyle o\_\{k\}=𝒞edit​\(Q,Ψ,Itret,ιk\),\\displaystyle=\\mathcal\{C\}\_\{\\mathrm\{edit\}\}\(Q,\\Psi,I\_\{t\}^\{\\mathrm\{ret\}\},\\iota\_\{k\}\),\(11\)ok∈\\displaystyle o\_\{k\}\\in\{Keep,Archive,Revise\},\\displaystyle\\\{\\text\{Keep\},\\text\{Archive\},\\text\{Revise\}\\\},whereΨ\\Psidenotes the final task feedback\. In addition, the controller separately decides whether the current task yields a new reusable insightκ^\\hat\{\\kappa\}:

a=𝒞add​\(Q,Ψ,Itret,κ^\),a∈\{0,1\}\.a=\\mathcal\{C\}\_\{\\mathrm\{add\}\}\(Q,\\Psi,I\_\{t\}^\{\\mathrm\{ret\}\},\\hat\{\\kappa\}\),\\quad a\\in\\\{0,1\\\}\.\(12\)This separation avoids conflating the correction of existing insights with the creation of new memory nodes\.

#### Archive\.

If an insight is judged to be invalid or harmful for the current task, we deactivate the node and record the current query as negative evidence:

\(κk,Ωk\+,Ωk−,zk\)←\(κk,Ωk\+,Ωk−∪\{q\},0\)\.\(\\kappa\_\{k\},\\Omega\_\{k\}^\{\+\},\\Omega\_\{k\}^\{\-\},z\_\{k\}\)\\leftarrow\(\\kappa\_\{k\},\\Omega\_\{k\}^\{\+\},\\Omega\_\{k\}^\{\-\}\\cup\\\{q\\\},0\)\.\(13\)The archived node is excluded from future retrieval but retained in the graph to preserve the failure evidence and support edit traceability\.

#### Revise\.

If an insight is partially useful but over\-generalized or outdated, the controller archives the old node and creates a revised node\. The old node is updated as

\(κk,Ωk\+,Ωk−,zk\)←\(κk,Ωk\+,Ωk−∪\{q\},0\)\.\(\\kappa\_\{k\},\\Omega\_\{k\}^\{\+\},\\Omega\_\{k\}^\{\-\},z\_\{k\}\)\\leftarrow\(\\kappa\_\{k\},\\Omega\_\{k\}^\{\+\},\\Omega\_\{k\}^\{\-\}\\cup\\\{q\\\},0\)\.\(14\)Then, a revised insight node is created:

ιk′=\(κk′,Ωk\+∪\{q\},∅,1\)\.\\iota\_\{k^\{\\prime\}\}=\(\\kappa^\{\\prime\}\_\{k\},\\Omega\_\{k\}^\{\+\}\\cup\\\{q\\\},\\varnothing,1\)\.\(15\)This operation makes the correction trace explicit: the old insight is preserved as an invalidated version, while the revised insight is activated as a corrected memory for future tasks\.

#### Add\.

If the current task reveals a reusable pattern that is not covered by existing retrieved insights, the controller creates a new insight node:

ιnew=\(κ^,\{q\},∅,1\)\.\\iota\_\{\\mathrm\{new\}\}=\(\\hat\{\\kappa\},\\\{q\\\},\\varnothing,1\)\.\(16\)Here,κ^\\hat\{\\kappa\}denotes the newly distilled insight and the current queryqqserves as its initial positive evidence\. This operation allows the memory graph to expand only when the current task contributes novel reusable knowledge\.

#### Keep\.

If a retrieved insight contributes positively to the current task, the controller preserves the insight and strengthens its positive evidence:

\(κk,Ωk\+,Ωk−,zk\)←\(κk,Ωk\+∪\{q\},Ωk−,1\)\.\(\\kappa\_\{k\},\\Omega\_\{k\}^\{\+\},\\Omega\_\{k\}^\{\-\},z\_\{k\}\)\\leftarrow\(\\kappa\_\{k\},\\Omega\_\{k\}^\{\+\}\\cup\\\{q\\\},\\Omega\_\{k\}^\{\-\},1\)\.\(17\)Repeated positive evidence increases the likelihood that the insight will be retrieved for future structurally related tasks\.

#### Edge Update\.

In addition to node\-level maintenance, the controller updates the insight graph connectivity\. When a new or revised insight is created, we connect it with the retrieved insights that jointly contributed to the current task:

ℰi←\\displaystyle\\mathcal\{E\}\_\{i\}\\leftarrowℰi∪\{\(ιa,ιb,q\)∣ιa∈Itret,\\displaystyle\\mathcal\{E\}\_\{i\}\\cup\\\{\(\\iota\_\{a\},\\iota\_\{b\},q\)\\mid\\iota\_\{a\}\\in I\_\{t\}^\{\\mathrm\{ret\}\},\(18\)ιb\\displaystyle\\iota\_\{b\}∈ℐnew,Ψ=Resolved\},\\displaystyle\\in\\mathcal\{I\}\_\{\\mathrm\{new\}\},\\Psi=\\text\{Resolved\}\\\},whereℐnew\\mathcal\{I\}\_\{\\mathrm\{new\}\}denotes the set of newly added or revised active insight nodes\. For archived nodes, existing edges are retained for traceability but ignored during active retrieval\. This update allows the graph to encode not only which insights are valid, but also how corrected insights emerge from prior memory contexts\.

Table 1:Performance comparison of different memory methods on PDDL, HotpotQA, and FEVER datasets\.![Refer to caption](https://arxiv.org/html/2608.11248v1/case.png)Figure 2:Case study demonstrating the evolution of insights upon FEVER tasks![Refer to caption](https://arxiv.org/html/2608.11248v1/ablation.png)Figure 3:Ablation results on the benchmarks with GPT\-4o\-mini and qwen2\.5\-7bTable 2:Average token consumption across PDDL, HotpotQA, and FEVER\.ΔNone\\Delta\_\{\\text\{None\}\}denotes the relative token change compared with the no\-memory baseline\.

## 4Experiments

### 4\.1Experiment Setup

#### Datasets\.

We conduct our experiments on PDDL\(Maet al\.,[2024](https://arxiv.org/html/2608.11248#bib.bib4)\), HotpotQAYanget al\.\([2018](https://arxiv.org/html/2608.11248#bib.bib6)\)and FEVERThorneet al\.\([2018](https://arxiv.org/html/2608.11248#bib.bib5)\)\.The PDDL dataset is a collection of strategic games—including Gripper, Barman, Blocksworld, and Tyreworld—adapted to evaluate agents in multi\-turn scenarios\. It requires multiple rounds of planned actions to finish a single subgoal, testing the agent’s ability to plan strategically and avoid repetitive steps\. The HotpotQA dataset challenges systems to find and reason over multiple supporting documents to arrive at an answer, providing sentence\-level supporting facts to facilitate explainable predictions\. The FEVER dataset is a large\-scale benchmark for fact extraction and verification\. It evaluates a model’s ability to retrieve necessary textual evidence and classify claims intoSupported,RefutedorNotEnoughInfocategories\.

#### Evaluation Metrics\.

For the FEVER and HotpotQA datasets, we useexact matchaccuracy\. For the PDDL datasets, we usethe progress rateto measure the progress of a task\.

#### Compared Methods\.

We compare ours with several representative memory methods, including:

#### MemoryBankZhonget al\.\([2024](https://arxiv.org/html/2608.11248#bib.bib7)\):

The framework is a long\-term memory mechanism tailored for Large Language Models\. It enables AI to store, recall, and update past interactions\. Inspired by the Ebbinghaus Forgetting Curve, it continually adapts to a user’s personality\.

#### VoyagerWanget al\.\([2023](https://arxiv.org/html/2608.11248#bib.bib8)\):

This method is the first LLM\-powered embodied lifelong learning agent in Minecraft that continuously explores the world, acquires diverse skills, and makes novel discoveries without human intervention\. It features an automatic curriculum, a skill library, and an iterative prompting mechanism\.

#### Generative AgentsParket al\.\([2023](https://arxiv.org/html/2608.11248#bib.bib9)\):

It operates on a system comprising foundational observational records and higher\-order cognitive reflections\. By deducing patterns and logically synthesizing fragmented interactions, the reflection layer distills abstract beliefs—ultimately elevating raw experiences into a highly structured and conceptually deep knowledge representation\.

#### G\-MemoryZhanget al\.\([2025](https://arxiv.org/html/2608.11248#bib.bib1)\):

This framework is a hierarchical graph memory architecture\. It manages lengthy collaboration histories through Insight, Query, and Interaction graph layers, enabling agents to efficiently retrieve past experiences for continuous team learning and self\-evolution\.

### 4\.2Main Results

The experimental results in PDDL, HotpotQA and FEVER datasets are shown in Table[1](https://arxiv.org/html/2608.11248#S3.T1)\. We summarize the key observations as follows:

\(1\) Among the evaluated memory methods, MemoryBank and Voyager demonstrate relatively weaker performance across the benchmarks\. This indicates that simply relying on the Ebbinghaus Forgetting Curve for memory decay or maintaining a fixed skill library is insufficient for dynamically managing complex, multi\-turn task contexts\. Generative Agents attempts to bridge this gap by synthesizing fragmented interactions into higher\-order cognitive reflections, however, its unstructured reflection mechanism struggles to maintain precise semantic dependencies, leading to suboptimal reasoning\. G\-Memory significantly outperforms these baselines by organizing lengthy histories into a structured, hierarchical graph architecture, yet it lacks a systematic memory management mechanism\. G\-Memory operates on a conventional append\-only design, this inevitably causes memory overload and introduces noise from non\-generalizable or obsolete insights\.

\(2\) Our failure\-aware memory maintenance framework resolves these issues by introducing a dynamic Graph Controller and an extended Insight Node structure\. Instead of passively accumulating data, our revised version supports active memory manipulation—such as modifying obsolete memories and pruning invalid representations\. To prevent interference of different queries, we partition the supporting queries into positive \(Ωk\+\\Omega\_\{k\}^\{\+\}\) and negative \(Ωk−\\Omega\_\{k\}^\{\-\}\) subsets , utilizing a conflict penalty mechanism to dynamically score and rank memory utility\. This ensures that only the most task\-aligned, highly generalizable insights are retrieved while mitigating the adverse impact of irrelevant data\. By explicitly addressing the structural flaws of prior methods, our approach yields a 10\.44% and 11\.38% progress rate improvement on the multi\-turn PDDL dataset, alongside a 21\.75% and 5\.03% exact match accuracy enhancement on the complex HotpotQA dataset across the two respective models\.

\(3\) We further analyze the token consumption of different memory methods in Table[2](https://arxiv.org/html/2608.11248#S3.T2)\. As expected, memory\-augmented methods generally consume more tokens than the no\-memory baseline, since they introduce additional retrieval, summarization, or memory\-update steps\. Among them, our method incurs the highest overall token usage, with a relative increase of 73\.0% on GPT\-4o\-mini and 72\.7% on Qwen2\.5\-7B compared with the no\-memory setting\. This overhead mainly comes from evidence\-aware insight retrieval and graph\-controller\-based memory correction\. However, compared with G\-Memory, which is the strongest graph\-based baseline, the additional token cost of our method is relatively moderate: token usage increases from 6\.2M to 6\.4M on GPT\-4o\-mini and from 5\.3M to 5\.7M on Qwen2\.5\-7B\. Considering the consistent performance gains reported in Table[1](https://arxiv.org/html/2608.11248#S3.T1), these results suggest that our framework trades a limited amount of additional computation for more reliable memory utilization\. These results indicate a performance\-cost trade\-off: while corrective memory maintenance introduces additional token overhead, it also improves the reliability and task utility of long\-term memory\.

### 4\.3Ablation Study

To evaluate the contribution of different operations in the proposed graph controller, we compare three variants:add\+keep, which only appends new insights and retains existing ones;add\+keep\+archive, which further allows the controller to deactivate outdated or harmful insights; andall, which uses the complete graph\-controller design\.

As shown in Figure[3](https://arxiv.org/html/2608.11248#S3.F3), the complete controller generally achieves the strongest performance across both backbone models\. On GPT\-4o\-mini, the full variant obtains 44%, 31%, and 68% on HotpotQA, PDDL, and FEVER, respectively\. Compared with the simpleadd\+keepvariant, this corresponds to gains of 4 points on HotpotQA, 5 points on PDDL, and 5 points on FEVER\. A similar trend can be observed on Qwen2\.5\-7B, where the full controller achieves 39%, 23%, and 65%, outperformingadd\+keepby 4, 4, and 4 points on the three benchmarks, respectively\.

The comparison betweenadd\+keepandadd\+keep\+archivefurther demonstrates the importance of memory deactivation\. Adding the archive operation improves performance in most settings, especially on GPT\-4o\-mini PDDL, where the score increases from 26% to 34%, and on FEVER, where it improves from 63% to 66%\. This suggests that long\-term memory should not be treated as a purely append\-only repository: obsolete or misleading insights can introduce noise into retrieval and harm downstream decision making\.

### 4\.4Case Study

To visually demonstrate whether the memory within the graph is effectively updated, we extract the system’s insight retrieval capabilities, alongside the explicit content of the recalled insights, across two consecutive tasks\. We present case studies in Figure[2](https://arxiv.org/html/2608.11248#S3.F2)\. As illustrated in the figure, upon receiving the initial task "Mohra is a truck", the system retrieves the insight "Verify claims by consulting multiple authoritative sources…" from memory to guide the ongoing task execution\. Subsequently, during the execution of the second task, the recalled insight transitions to "Verify claims by cross\-referencing with multiple authoritative sources and considering the context…"\. This progression demonstrates that the insights undergo dynamic self\-evolution and refinement throughout the sequence of tasks\. Consequently, these adjusted insights become increasingly fine\-grained and better adapted to novel requirements, playing a pivotal role in facilitating ultimate task success\.

## 5Conclusion

In this paper, we studied insight\-level memory maintenance for long\-term language agents, where previously distilled insights may become ineffective, over\-generalized, or harmful under new task contexts\. We proposed a failure\-aware memory maintenance framework that extends graph\-based agent memory with editable insight nodes\. Each insight tracks positive evidence, negative evidence, and an activation state, while a graph controller updates the memory graph through keeping, archiving, revising, and adding insights after task execution\.

Experiments on PDDL, HotpotQA, and FEVER show that our method improves over representative memory\-based agent frameworks across different backbone models\. The ablation study further suggests that append\-only memory is insufficient for long\-horizon tasks, and that evidence\-aware retrieval together with graph\-level editing contributes to stronger performance\. Overall, our work shifts long\-term agent memory from passive accumulation toward corrective maintenance\.

## 6Limitations

Our framework has several limitations\. First, the graph controller depends on post\-task feedback, and noisy or incomplete feedback may lead to suboptimal edits, such as archiving useful insights or revising memories prematurely\. Second, the positive and negative evidence sets provide a simple validity signal but do not fully capture the context or degree of an insight’s usefulness\. Third, our experiments focus on PDDL, HotpotQA, and FEVER, leaving broader long\-term agent scenarios, such as web interaction, software engineering, and personalized assistants, for future evaluation\. Finally, maintaining and editing an insight graph introduces additional computational and storage overhead, which may require more efficient pruning and controller invocation strategies in large\-scale deployments\.

## References

- P\. Anokhin, N\. Semenov, A\. Sorokin, D\. Evseev, A\. Kravchenko, M\. Burtsev, and E\. Burnaev \(2024\)Arigraph: learning knowledge graph world models with episodic memory for llm agents\.arXiv preprint arXiv:2407\.04363\.Cited by:[§2\.2](https://arxiv.org/html/2608.11248#S2.SS2.p1.1)\.
- P\. Chhikara, D\. Khant, S\. Aryan, T\. Singh, and D\. Yadav \(2025\)Mem0: building production\-ready ai agents with scalable long\-term memory\.arXiv preprint arXiv:2504\.19413\.Cited by:[§2\.1](https://arxiv.org/html/2608.11248#S2.SS1.p1.1)\.
- J\. Fang, H\. Jiang, K\. Wang, Y\. Ma, J\. Shi, X\. Wang, X\. He, and T\. Chua \(2025\)Alphaedit: null\-space constrained knowledge editing for language models\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 16366–16396\.Cited by:[§2\.2](https://arxiv.org/html/2608.11248#S2.SS2.p2.1)\.
- B\. J\. Gutiérrez, Y\. Shu, Y\. Gu, M\. Yasunaga, and Y\. Su \(2024\)Hipporag: neurobiologically inspired long\-term memory for large language models\.Advances in neural information processing systems37,pp\. 59532–59569\.Cited by:[§2\.2](https://arxiv.org/html/2608.11248#S2.SS2.p1.1)\.
- T\. Hartvigsen, S\. Sankaranarayanan, H\. Palangi, Y\. Kim, and M\. Ghassemi \(2023\)Aging with grace: lifelong model editing with discrete key\-value adaptors\.Advances in Neural Information Processing Systems36,pp\. 47934–47959\.Cited by:[§2\.2](https://arxiv.org/html/2608.11248#S2.SS2.p2.1)\.
- J\. Kang, M\. Ji, Z\. Zhao, and T\. Bai \(2025\)Memory os of ai agent\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 25972–25981\.Cited by:[§2\.1](https://arxiv.org/html/2608.11248#S2.SS1.p1.1)\.
- Q\. Li and X\. Chu \(2025\)AdaEdit: advancing continuous knowledge editing for large language models\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 4127–4149\.External Links:[Link](https://aclanthology.org/2025.acl-long.208/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.208),ISBN 979\-8\-89176\-251\-0Cited by:[§2\.2](https://arxiv.org/html/2608.11248#S2.SS2.p2.1)\.
- C\. Ma, J\. Zhang, Z\. Zhu, C\. Yang, Y\. Yang, Y\. Jin, Z\. Lan, L\. Kong, and J\. He \(2024\)Agentboard: an analytical evaluation board of multi\-turn llm agents\.Advances in neural information processing systems37,pp\. 74325–74362\.Cited by:[§4\.1](https://arxiv.org/html/2608.11248#S4.SS1.SSS0.Px1.p1.1)\.
- K\. Meng, D\. Bau, A\. Andonian, and Y\. Belinkov \(2022a\)Locating and editing factual associations in gpt\.Advances in neural information processing systems35,pp\. 17359–17372\.Cited by:[§2\.2](https://arxiv.org/html/2608.11248#S2.SS2.p2.1)\.
- K\. Meng, A\. S\. Sharma, A\. Andonian, Y\. Belinkov, and D\. Bau \(2022b\)Mass\-editing memory in a transformer\.arXiv preprint arXiv:2210\.07229\.Cited by:[§2\.2](https://arxiv.org/html/2608.11248#S2.SS2.p2.1)\.
- E\. Mitchell, C\. Lin, A\. Bosselut, C\. Finn, and C\. D\. Manning \(2021\)Fast model editing at scale\.arXiv preprint arXiv:2110\.11309\.Cited by:[§1](https://arxiv.org/html/2608.11248#S1.p3.1),[§2\.2](https://arxiv.org/html/2608.11248#S2.SS2.p2.1)\.
- C\. Packer, V\. Fang, S\. Patil, K\. Lin, S\. Wooders, and J\. Gonzalez \(2023\)MemGPT: towards llms as operating systems\.\.Cited by:[§2\.1](https://arxiv.org/html/2608.11248#S2.SS1.p1.1)\.
- J\. S\. Park, J\. O’Brien, C\. J\. Cai, M\. R\. Morris, P\. Liang, and M\. S\. Bernstein \(2023\)Generative agents: interactive simulacra of human behavior\.InProceedings of the 36th annual acm symposium on user interface software and technology,pp\. 1–22\.Cited by:[§2\.1](https://arxiv.org/html/2608.11248#S2.SS1.p1.1),[§4\.1](https://arxiv.org/html/2608.11248#S4.SS1.SSS0.Px6)\.
- P\. Rasmussen, P\. Paliychuk, T\. Beauvais, J\. Ryan, and D\. Chalef \(2025\)Zep: a temporal knowledge graph architecture for agent memory\.arXiv preprint arXiv:2501\.13956\.Cited by:[§2\.2](https://arxiv.org/html/2608.11248#S2.SS2.p1.1)\.
- N\. Shinn, F\. Cassano, A\. Gopinath, K\. Narasimhan, and S\. Yao \(2023\)Reflexion: language agents with verbal reinforcement learning\.Advances in neural information processing systems36,pp\. 8634–8652\.Cited by:[§2\.1](https://arxiv.org/html/2608.11248#S2.SS1.p1.1)\.
- J\. Thorne, A\. Vlachos, C\. Christodoulopoulos, and A\. Mittal \(2018\)FEVER: a large\-scale dataset for fact extraction and verification\.InProceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long Papers\),pp\. 809–819\.Cited by:[§4\.1](https://arxiv.org/html/2608.11248#S4.SS1.SSS0.Px1.p1.1)\.
- G\. Wang, Y\. Xie, Y\. Jiang, A\. Mandlekar, C\. Xiao, Y\. Zhu, L\. Fan, and A\. Anandkumar \(2023\)Voyager: an open\-ended embodied agent with large language models\.arXiv preprint arXiv:2305\.16291\.Cited by:[§1](https://arxiv.org/html/2608.11248#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.11248#S2.SS1.p1.1),[§4\.1](https://arxiv.org/html/2608.11248#S4.SS1.SSS0.Px5)\.
- W\. Xu, Z\. Liang, K\. Mei, H\. Gao, J\. Tan, and Y\. Zhang \(2026\)A\-mem: agentic memory for llm agents\.Advances in Neural Information Processing Systems38,pp\. 17577–17604\.Cited by:[§2\.1](https://arxiv.org/html/2608.11248#S2.SS1.p1.1)\.
- Z\. Yang, P\. Qi, S\. Zhang, Y\. Bengio, W\. Cohen, R\. Salakhutdinov, and C\. D\. Manning \(2018\)HotpotQA: a dataset for diverse, explainable multi\-hop question answering\.InProceedings of the 2018 conference on empirical methods in natural language processing,pp\. 2369–2380\.Cited by:[§4\.1](https://arxiv.org/html/2608.11248#S4.SS1.SSS0.Px1.p1.1)\.
- G\. Zhang, M\. Fu, K\. Wang, G\. Wan, M\. Yu, and S\. YAN \(2025\)G\-memory: tracing hierarchical memory for multi\-agent systems\.InThe Thirty\-ninth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=mmIAp3cVS0)Cited by:[§1](https://arxiv.org/html/2608.11248#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.11248#S2.SS2.p1.1),[§3](https://arxiv.org/html/2608.11248#S3.p1.1),[§4\.1](https://arxiv.org/html/2608.11248#S4.SS1.SSS0.Px7)\.
- W\. Zhong, L\. Guo, Q\. Gao, H\. Ye, and Y\. Wang \(2024\)Memorybank: enhancing large language models with long\-term memory\.InProceedings of the AAAI conference on artificial intelligence,Vol\.38,pp\. 19724–19731\.Cited by:[§1](https://arxiv.org/html/2608.11248#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.11248#S2.SS1.p1.1),[§4\.1](https://arxiv.org/html/2608.11248#S4.SS1.SSS0.Px4)\.

## Appendix Appendix Aprompt\-constrained Graph Controller

This section presents the prompt used to instantiate the prompt\-constrained graph controller in our memory maintenance framework\. Given the current query, the agent trajectory, the final answer, task feedback, and the retrieved insight nodes, the controller is required to output structured memory edit decisions\. Specifically, it selects one operation fromKeep,Archive, andRevisefor each retrieved insight, and separately determines whether a new reusable insight should be added to the memory graph\. The output is constrained to a predefined JSON format, which enables stable and automatic updates of positive evidence, negative evidence, and activation states in the editable insight graph\.

Graph ControllerYou are a graph memory controller\. Update retrieved insight nodes after a task\.For each retrieved insight, choose exactly one operation:KEEP: Choose this if the insight was valid and directly useful for the current task\. The current query will be added to positive evidence\.ARCHIVE: Choose this if the insight was invalid, harmful, misleading, or contradicted by the current task\. The current query will be added to negative evidence and the node will be deactivated\.REVISE: Choose this if the insight was partially useful but too broad, outdated, or missing necessary conditions\. The old node will be archived, and a new revised active node will be created\. The revised insight must correct or narrow the original insight, not merely paraphrase it\.ADD: After editing existing insights, decide whether the current task contains a new reusable insight not covered by retrieved insights\. Add only generalizable and non\-duplicated insights\.Use only the given task information\. Be conservative\. Do not create task\-specific memories\.Current query:\{\{CURRENT\_QUERY\}\}Trajectory:\{\{TRAJECTORY\}\}Final answer:\{\{FINAL\_ANSWER\}\}Task feedback:\{\{TASK\_FEEDBACK\}\}Retrieved insights:\{\{RETRIEVED\_INSIGHTS\_JSON\}\}Return only valid JSON in this format:\{"insight\_edits": \[\{"insight\_id": "\.\.\.","operation": "KEEP\|ARCHIVE\|REVISE","revised\_insight\_text": null\}\],"add\_new\_insight": \{"decision": false,"new\_insight\_text": null\}\}

Similar Articles

EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments

Hugging Face Daily Papers

EvoArena introduces a benchmark for evaluating LLM agents in dynamic environments with progressive updates across terminal, software, and social domains, while EvoMem proposes a patch-based memory paradigm that records structured evolution; experiments show current agents achieve only 39.6% accuracy on EvoArena, and EvoMem yields average gains of 1.5% on the benchmark and improvements on GAIA and LoCoMo.

G-Long: Graph-Enhanced Memory Management for Efficient Long-Term Dialogue Agents

arXiv cs.CL

G-Long proposes a graph-enhanced memory management framework for long-term dialogue agents, using a fine-tuned small language model for structured triplet extraction and associative retrieval, achieving state-of-the-art performance in response generation and memory retrieval with reduced computational overhead.

MemEvoBench: Benchmarking Memory MisEvolution in LLM Agents

arXiv cs.CL

MemEvoBench introduces the first benchmark for evaluating memory safety in LLM agents, measuring behavioral degradation from adversarial memory injection, noisy outputs, and biased feedback across QA and workflow tasks. The work reveals that memory evolution significantly contributes to safety failures and that static defenses are insufficient.