From Faulty Memories to Corrected Actions: Dependency-Guided Rollback Repair for Memory-Augmented Agents

arXiv cs.AI Papers

Summary

This paper introduces dependency-guided rollback repair for memory-augmented agents, a method that builds a typed memory-to-action graph from runtime provenance to selectively undo faulty memory effects while preserving benign state, achieving strong recovery on benchmarks.

arXiv:2608.10502v1 Announce Type: new Abstract: Persistent memory lets language-model agents reuse information across sessions, but it also makes errors durable: a poisoned, stale, or misattributed record can alter reasoning, tool use, answers, and subsequent memory writes. Existing defenses mainly detect or delete suspicious memories, or revise the current response. Deleting the source leaves already propagated claims, actions, and derived memories active, whereas resetting the store or replaying the full trace destroys benign state and repeats unnecessary computation. We therefore formulate \textbf{post-failure memory recovery: } \textit{given a failed execution and diagnosed faulty memories, recover both the answer and persistent state while retaining unaffected work.} Our \textbf{dependency-guided rollback repair} builds a typed memory-to-action graph from runtime provenance, traces explicit downstream dependencies, preserves candidates with independent trusted support, deactivates unsupported memory state, and selectively replays only answer-relevant affected computation. We evaluate this approach on a 150-case controlled benchmark spanning three tool-use domains and four memory failure types, and on a 50-case trajectory-derived stress test adapted from LongMemEval-V2. On the controlled benchmark, it achieves 85.3\% recovery versus 77.3\% for the best competing recovery method, removes all diagnosed faulty memories, preserves all benign memories, and requires only selective replay with modest LLM-call cost. On the adapted subset, it reaches 68.0\% recovery versus 54.0\% for the next best method, while also achieving the highest claim invalidation F1, 0.669 versus 0.603. Overall, the results do not imply uniformly better trace reconstruction, but show that dependency-guided rollback repair provides a strong recovery--cost trade-off while repairing faulty memory state and preserving benign memory.
Original Article
View Cached Full Text

Cached at: 08/12/26, 08:25 AM

# From Faulty Memories to Corrected Actions: Dependency-Guided Rollback Repair for Memory-Augmented Agents
Source: [https://arxiv.org/html/2608.10502](https://arxiv.org/html/2608.10502)
Caili Yu, Yiqi Wang\*, Jiaqi Zhang, Yiqun Duan, Mingkai Zheng, Zhangkai Wu, Kaize Shi, Taotao Cai \*yiqi\.wang\.jennie@gmail\.com

###### Abstract

Persistent memory lets language\-model agents reuse information across sessions, but it also makes errors durable: a poisoned, stale, or misattributed record can alter reasoning, tool use, answers, and subsequent memory writes\. Existing defenses mainly detect or delete suspicious memories, or revise the current response\. Deleting the source leaves already propagated claims, actions, and derived memories active, whereas resetting the store or replaying the full trace destroys benign state and repeats unnecessary computation\. We therefore formulatepost\-failure memory recovery:given a failed execution and diagnosed faulty memories, recover both the answer and persistent state while retaining unaffected work\.Ourdependency\-guided rollback repairbuilds a typed memory\-to\-action graph from runtime provenance, traces explicit downstream dependencies, preserves candidates with independent trusted support, deactivates unsupported memory state, and selectively replays only answer\-relevant affected computation\. We evaluate this approach on a 150\-case controlled benchmark spanning three tool\-use domains and four memory failure types, and on a 50\-case trajectory\-derived stress test adapted from LongMemEval\-V2\. On the controlled benchmark, it achieves 85\.3% recovery versus 77\.3% for the best competing recovery method, removes all diagnosed faulty memories, preserves all benign memories, and requires only selective replay with modest LLM\-call cost\. On the adapted subset, it reaches 68\.0% recovery versus 54\.0% for the next best method, while also achieving the highest claim invalidation F1, 0\.669 versus 0\.603\. Overall, the results do not imply uniformly better trace reconstruction, but show that dependency\-guided rollback repair provides a strong recovery–cost trade\-off while repairing faulty memory state and preserving benign memory\.

## 1Introduction

Persistent memory lets language model agents carry preferences, observations, and experience across sessions, enabling personalization and long\-horizon behavior that a current prompt alone cannot support\(Parket al\.[2023](https://arxiv.org/html/2608.10502#bib.bib1); Packeret al\.[2023](https://arxiv.org/html/2608.10502#bib.bib2); Zhanget al\.[2025](https://arxiv.org/html/2608.10502#bib.bib3); Zhonget al\.[2024](https://arxiv.org/html/2608.10502#bib.bib4); Wanget al\.[2023](https://arxiv.org/html/2608.10502#bib.bib5)\)\. The same persistence changes the scope of failure\. A poisoned, stale, misattributed, or drifted record can be retrieved as context, support an incorrect claim, alter a tool plan and answer, and be consolidated into new memory\(Xionget al\.[2026](https://arxiv.org/html/2608.10502#bib.bib6); Chaoet al\.[2026](https://arxiv.org/html/2608.10502#bib.bib7); Chenet al\.[2025](https://arxiv.org/html/2608.10502#bib.bib8)\)\. A single local fault can therefore become a durable family of downstream errors and reappear after the original turn\.

This propagation creates the practical recovery problem illustrated in Figure[1](https://arxiv.org/html/2608.10502#S1.F1)\. Correcting only the answer leaves the faulty source and contaminated descendants available to influence future turns\. Deleting only the source is also insufficient after its content has been copied into claims, observations, or derived memories\. At the other extreme, clearing the memory store or replaying the full trace removes valid personalization and repeats model and tool calls\. Useful recovery must therefore remove unsupported consequences, retain state that still has independent evidence, and recompute only what the corrected answer requires\.

![Refer to caption](https://arxiv.org/html/2608.10502v1/x1.png)Figure 1:A faulty memory creates downstream state\.A poisoned preference changes the claim, tool plan, answer, and a derived memory\. Repair must remove unsupported descendants, preserve unaffected state, and replay the answer\-relevant path\.Existing research largely acts at one of two boundaries\. Memory defenses audit, filter, or delete suspicious records, while response\-level reflection revises an unsuccessful output\(Tanet al\.[2026](https://arxiv.org/html/2608.10502#bib.bib10); Zouet al\.[2025](https://arxiv.org/html/2608.10502#bib.bib11); Shinnet al\.[2023](https://arxiv.org/html/2608.10502#bib.bib9)\)\. These interventions improve detection or local correction, but do not jointly repair the persistent store and the execution state already influenced by a fault\. This gap motivates our central premise:*fault removal is not state recovery*\.

We therefore formulatepost\-failure memory recovery\. The input is a user session, an active memory store, a failed execution trace, and diagnosed faulty memories supplied by an upstream detector; the output is a corrected answer, repaired trace, and repaired active store\. Diagnosis itself is outside the scope of this work, and the runtime must record the provenance connecting memory reads, claims, plans, actions, observations, answers, and mutations\. Under these assumptions, the research question is precise:*once a faulty memory has been consumed, which consequences should be invalidated, which remain independently justified, and which must be regenerated?*

We propose adependency\-guided rollback repair\. The method first constructs a typed memory\-to\-action graph and traces explicit downstream dataflow from each diagnosed fault\. Because reachability alone would over\-invalidate, it preserves candidates whose complete content or decision has independent trusted support\. A deterministic planner then deletes diagnosed faults, quarantines unsupported derived state, invalidates unsupported trace outputs, and selects the answer\-relevant affected computation\. Finally, the executor reuses safe context and replays selected steps in trace order under the repaired store\. The design directly couples the three requirements above: contamination tracing, evidence\-aware preservation, and selective recomputation\.

We evaluate answer recovery, state repair, preservation, and cost on 150 controlled cases across shopping, travel, and customer\-support tools, covering four memory fault types, and on 50 procedural trajectories adapted from LongMemEval\-V2 as a multi\-fault stress test\(Wuet al\.[2026](https://arxiv.org/html/2608.10502#bib.bib35)\)\. Our method obtains the highest immediate recovery in both settings, although it does not achieve the lowest recurrence on the controlled benchmark\. Overall, the results do not imply uniformly better trace reconstruction, but show that dependency\-guided rollback repair provides a strong balance among recovery, repair cost, faulty\-state cleanup, and benign\-state preservation\.

Our main contributions are summarized as follows:

- •We formulatepost\-failure memory recovery, which jointly considers answer recovery, downstream state repair, benign\-state preservation, and repair cost after faulty memories have already affected agent execution\.
- •We propose adependency\-guided rollback repairmethod combining explicit dataflow tracing, independent\-support checking, rule\-guided planning, and answer\-relevant replay\.
- •We introduce a150\-case controlled benchmarkplus a 50\-case trajectory\-derived stress test, showing the highest immediate recovery and a favorable recovery–cost trade\-off\.

## 2Related Work

#### Agent memory and failure detection\.

Persistent\-memory systems support cross\-session personalization and experience reuse\(Parket al\.[2023](https://arxiv.org/html/2608.10502#bib.bib1); Packeret al\.[2023](https://arxiv.org/html/2608.10502#bib.bib2); Xuet al\.[2026](https://arxiv.org/html/2608.10502#bib.bib13); Duet al\.[2026](https://arxiv.org/html/2608.10502#bib.bib15); Zhaoet al\.[2026b](https://arxiv.org/html/2608.10502#bib.bib14)\)\. Recent work studies error propagation, hallucinated updates, stale state, retrieval or memory poisoning, and post\-hoc auditing\(Xionget al\.[2026](https://arxiv.org/html/2608.10502#bib.bib6); Chenet al\.[2025](https://arxiv.org/html/2608.10502#bib.bib8); Chaoet al\.[2026](https://arxiv.org/html/2608.10502#bib.bib7); Zouet al\.[2025](https://arxiv.org/html/2608.10502#bib.bib11); Chenet al\.[2024](https://arxiv.org/html/2608.10502#bib.bib32); Tanet al\.[2026](https://arxiv.org/html/2608.10502#bib.bib10)\)\. These methods primarily improve what enters, remains in, or is retrieved from memory\. We instead condition on diagnosed faults that have already influenced an execution\.

#### Memory cascade repair\.

MemoRepair is the closest memory\-side recovery framework: it withdraws invalidated descendants, rebuilds successors from retained support, and uses predecessor closure for cost\-aware republication\(Zhaoet al\.[2026a](https://arxiv.org/html/2608.10502#bib.bib36)\)\. Our setting is complementary but broader in execution scope\. It jointly repairs persistent memories and a heterogeneous agent trace containing claims, plans, tool actions, observations, and the answer; answer relevance determines which invalidated execution nodes are replayed\. We do not claim to replace MemoRepair’s publication contract or fault detector\. Rather, we study end\-to\-end answer\-and\-state recovery once a diagnosed memory fault has crossed into agent execution\.

#### Provenance, slicing, and rollback\.

Forward dependency tracing and selective recomputation draw on program and dynamic slicing\(Weiser[1984](https://arxiv.org/html/2608.10502#bib.bib37); Agrawal and Horgan[1990](https://arxiv.org/html/2608.10502#bib.bib38)\), database provenance\(Cheneyet al\.[2009](https://arxiv.org/html/2608.10502#bib.bib39)\), and transactional rollback\(Mohanet al\.[1992](https://arxiv.org/html/2608.10502#bib.bib40)\)\. We do not claim graph reachability or logical undo as new\. Our contribution is their adaptation to a typed memory–execution graph in which validity depends on independent evidentiary support and replay incurs model and tool cost\. This scope is narrower than a new general recovery theory, but broader than repairing a memory record or final response alone\.

#### Trace debugging and response repair\.

ReAct exposes reasoning and tool\-use structure, while Reflexion and Self\-Refine revise unsuccessful behavior or outputs\(Yaoet al\.[2022](https://arxiv.org/html/2608.10502#bib.bib20); Shinnet al\.[2023](https://arxiv.org/html/2608.10502#bib.bib9); Madaanet al\.[2023](https://arxiv.org/html/2608.10502#bib.bib33)\)\. AgentTrace reconstructs causal graphs from logs for root\-cause localization\(Wang[2026](https://arxiv.org/html/2608.10502#bib.bib34)\)\. These ideas motivate trace\-aware recovery, but localization or answer revision alone need not deactivate contaminated persistent state\. Our dependency graph therefore spans the memory lifecycle and execution trace, and its rollback plan couples memory disposition with selective replay\.

## 3Dependency\-Guided Rollback Repair

![Refer to caption](https://arxiv.org/html/2608.10502v1/x2.png)Figure 2:Overview of dependency\-guided rollback repair\.Starting from a diagnosed faulty memory set, the method constructs a memory\-to\-action dependency graph, traces candidate downstream effects, filters candidates with independent trusted support, generates a rule\-guided rollback plan, and selectively replays answer\-relevant affected computation to produce a corrected final answer and a repaired active memory store\.As Figure[2](https://arxiv.org/html/2608.10502#S3.F2)shows, we propose dependency\-guided rollback repair as our method\. It proceeds in five stages: graph construction, affected\-subgraph tracing, independent\-support checking, rule\-guided rollback planning, and selective replay\. Rollback here means repair of agent\-maintained trace and memory state\. It cannot undo an irreversible external side effect\. Re\-executing a side\-effecting tool therefore requires a resettable interface or a domain\-specific compensating action\. Full schemas are in Appendix[A](https://arxiv.org/html/2608.10502#A1)and replay prompt contracts in Appendix[B](https://arxiv.org/html/2608.10502#A2)\.

### 3\.1Problem Setting

Let𝒰=\(u1,…,uk\)\\mathcal\{U\}=\(u\_\{1\},\\ldots,u\_\{k\}\)be a user session,ℳ\\mathcal\{M\}the active memory store, and𝒯=\(s1,…,sn\)\\mathcal\{T\}=\(s\_\{1\},\\ldots,s\_\{n\}\)a failed execution trace\. A step is a memory read, claim, plan, tool action, tool observation, answer, or memory mutation\. The input also contains diagnosed faulty memoriesℱ⊆ℳ\\mathcal\{F\}\\subseteq\\mathcal\{M\}\. At repair time, our method sees only\(𝒰,ℳ,𝒯,ℱ\)\(\\mathcal\{U\},\\mathcal\{M\},\\mathcal\{T\},\\mathcal\{F\}\), available tools, and runtime provenance; clean traces and evaluation labels are withheld\.

### 3\.2Dependency Graph Construction

We instantiate a directed heterogeneous graph𝒢=\(𝒱,ℰ\)\\mathcal\{G\}=\(\\mathcal\{V\},\\mathcal\{E\}\)from user inputs𝒰\\mathcal\{U\}, execution steps𝒯\\mathcal\{T\}, and memory recordsℳ\\mathcal\{M\}\. Construction is fault\-agnostic: it records execution dataflow and memory\-lifecycle structure without usingℱ\\mathcal\{F\}or any benchmark annotation\. The node set is partitioned as𝒱=𝒱𝒰∪𝒱𝒯∪𝒱ℳ\\mathcal\{V\}=\\mathcal\{V\}\_\{\\mathcal\{U\}\}\\cup\\mathcal\{V\}\_\{\\mathcal\{T\}\}\\cup\\mathcal\{V\}\_\{\\mathcal\{M\}\}, containing user inputs, typed execution steps, and persistent memory records; we identify each memory record with its node in𝒱ℳ\\mathcal\{V\}\_\{\\mathcal\{M\}\}when using memory sets in graph expressions\.

We instantiate edges deterministically from these fields rather than inferring them\. Edges point from a prerequisite or source to its dependent\. They encodeinitiate,cite,support,produce,delete,update,consolidaterelations; and memorysupersede,deriverelations\. Temporalchainedges retain execution order\. A node pair may carry multiple labels\. Appendix[A](https://arxiv.org/html/2608.10502#A1)gives the full edge semantics\.

Each trace record stores the identifiers of the memories, claims, observations, together with tool\-call identifiers and memory\-lifecycle metadata\. This provenance is emitted by the agent runtime, so the method presupposes an instrumented agent\. Inferring missing support or lifecycle edges from raw natural\-language logs is outside the scope of this work; incomplete provenance can under\-trace contamination, while spurious edges can cause unnecessary invalidation\.

For propagation we use

ℰprop=\\displaystyle\\mathcal\{E\}\_\{\\mathrm\{prop\}\}=\{\}ℰinitiate∪ℰcite∪ℰsupport∪ℰproduce\\displaystyle\\mathcal\{E\}\_\{\\mathrm\{initiate\}\}\\cup\\mathcal\{E\}\_\{\\mathrm\{cite\}\}\\cup\\mathcal\{E\}\_\{\\mathrm\{support\}\}\\cup\\mathcal\{E\}\_\{\\mathrm\{produce\}\}∪ℰdelete∪ℰupdate∪ℰconsolidate\\displaystyle\{\}\\cup\\mathcal\{E\}\_\{\\mathrm\{delete\}\}\\cup\\mathcal\{E\}\_\{\\mathrm\{update\}\}\\cup\\mathcal\{E\}\_\{\\mathrm\{consolidate\}\}∪ℰsupersede∪ℰderive\.\\displaystyle\{\}\\cup\\mathcal\{E\}\_\{\\mathrm\{supersede\}\}\\cup\\mathcal\{E\}\_\{\\mathrm\{derive\}\}\.Purechainedges are excluded: occurring later is not evidence of contamination\. This distinction is important because a sequential suffix may contain unrelated valid work\.

### 3\.3Fault Provenance & Affected Subgraph Tracing

Givenℱ\\mathcal\{F\}, we identify the state that may depend on these faults\. This stage does not use fault\-type labels such as poisoned, stale, wrong\-user, or summary\-drift; all diagnosed faults are handled uniformly through graph structure and memory provenance\.

For eachm∈ℱm\\in\\mathcal\{F\}, we use fault provenance to identify a reachability seedr​\(m\)r\(m\)\. We inspectdelete,update, andconsolidaterelations to recover lineage and identify mutations\. If a failed mutation exists, it becomes the seed; otherwise, we follow aproducerelation to the producing step\. Ifmmwas not created during the recorded execution and has no such producing step, the memory itself becomes the seed\. Whenr​\(m\)r\(m\)occurs at timetrt\_\{r\}in the trace, only dependents timestamped no earlier thantrt\_\{r\}are considered\.

We then trace the affected subgraph from eachr​\(m\)r\(m\)and take the union of the reached nodes and all seeds to form the raw affected candidate set:

𝒜0=⋃m∈ℱReach\(𝒱,ℰprop\)​\(r​\(m\)\)\.\\mathcal\{A\}\_\{0\}=\\bigcup\_\{m\\in\\mathcal\{F\}\}\\mathrm\{Reach\}\_\{\(\\mathcal\{V\},\\mathcal\{E\}\_\{\\mathrm\{prop\}\}\)\}\\bigl\(r\(m\)\\bigr\)\.A node is not marked as affected merely because it follows a faulty memory in execution order; it must be connected through an explicit propagation dependency\. Favoring recall,𝒜0\\mathcal\{A\}\_\{0\}over\-approximates the truly unsupported state and is refined next\.

### 3\.4Independent\-Support Checking

A node may be reachable from a faulty memory while remaining valid because it is also justified by independent, fault\-free evidence\. For example, a reasoning step may consume a faulty memory together with an explicit current\-turn instruction or a successful tool observation that independently supports the same conclusion\. Treating every reachable node as contaminated would therefore cause unnecessary invalidation and replay\.

For each semantically checkable candidatev∈𝒜0v\\in\\mathcal\{A\}\_\{0\}we collect its sufficient evidentiary predecessorsPredsup​\(v\)\\mathrm\{Pred\}\_\{\\mathrm\{sup\}\}\(v\)throughsupport,derive,supersederelations andciterelations marked sufficient by the provenance contract\. Under our fixed trust policy an admissible source is an explicit current\-turn instruction, an active memory not diagnosed as faulty, a successful tool observation, or an unaffected validated claim, and a source counts as independent only outsideℱ∪𝒜0\\mathcal\{F\}\\cup\\mathcal\{A\}\_\{0\}:TrustedSupport​\(v\)⇔∃p∈Predsup​\(v\)\\mathrm\{TrustedSupport\}\(v\)\\iff\\exists\\,p\\in\\mathrm\{Pred\}\_\{\\mathrm\{sup\}\}\(v\)withp∉ℱ∪𝒜0p\\notin\\mathcal\{F\}\\cup\\mathcal\{A\}\_\{0\}andAdmissible​\(p\)\\mathrm\{Admissible\}\(p\)\.

Diagnosed faulty memories are never eligible for preservation\. Claims, plans, answers, mutations, and active derived memories are checked directly; a derived memory requires at least one sufficient provenance path terminating at it whose source and internal supporting nodes are admissible and outsideℱ∪𝒜0\\mathcal\{F\}\\cup\\mathcal\{A\}\_\{0\}\. Actions inherit the verdicts of their plans, and observations inherit those of their generating actions, so an independently supported plan may keep its unchanged action and observation\. The check is conservative: only evidence outside𝒜0\\mathcal\{A\}\_\{0\}may preserve a candidate, and preserved candidates are not recursively reused to validate others, preventing mutually dependent affected nodes from validating one another\. Writing𝒫sup⊆𝒜0\\mathcal\{P\}\_\{\\mathrm\{sup\}\}\\subseteq\\mathcal\{A\}\_\{0\}for the preserved candidates, including associated safe action and observation nodes, the unsupported affected set is𝒜=𝒜0∖𝒫sup\\mathcal\{A\}=\\mathcal\{A\}\_\{0\}\\setminus\\mathcal\{P\}\_\{\\mathrm\{sup\}\}:𝒫sup\\mathcal\{P\}\_\{\\mathrm\{sup\}\}stays eligible for reuse and𝒜\\mathcal\{A\}passes to the planner\.

### 3\.5Rule\-Guided Rollback Planning

The rollback planner converts the diagnosed faults, unsupported affected set, provenance information, and support verdicts into the repair planℛ\\mathcal\{R\}\. The planner is deterministic and rule\-guided rather than optimization\-based\. Its memory dispositions and trace actions are orthogonal: deleting or quarantining a memory changes the active memory state, whereas invalidating, replaying, or preserving a trace node determines how the previous execution is treated\.

The planner first deactivates unsafe persistent state, settingℳdel=ℱ\\mathcal\{M\}\_\{\\mathrm\{del\}\}=\\mathcal\{F\}andℳquar=\(𝒜∩𝒱mem\)∖ℱ\\mathcal\{M\}\_\{\\mathrm\{quar\}\}=\(\\mathcal\{A\}\\cap\\mathcal\{V\}\_\{\\mathrm\{mem\}\}\)\\setminus\\mathcal\{F\}, so faulty memories leave the active store, unsupported affected memories are quarantined, and support\-preserved memories stay active; when a faulty memory has a recoverable producing step, or stale state results from a failed lifecycle operation, that step stays eligible for replay\. All unsupported trace nodes are then marked invalid,𝒱inv=𝒜∩𝒱step\\mathcal\{V\}\_\{\\mathrm\{inv\}\}=\\mathcal\{A\}\\cap\\mathcal\{V\}\_\{\\mathrm\{step\}\}, so their outputs cannot serve as evidence; plans pass invalidation to their actions, and actions pass it to their observations, and observations are never replayed directly, only refreshed by scheduling the generating action\. Appendix[C](https://arxiv.org/html/2608.10502#A3)tabulates the full rule set\.

To avoid replaying the full affected subgraph, the planner next identifies the answer\-relevant affected computation\. Letvfinalv\_\{\\mathrm\{final\}\}denote the final\-answer node, and let

𝒢−=𝒢dep​\[𝒱∖\(ℳdel∪ℳquar\)\]\\mathcal\{G\}^\{\-\}=\\mathcal\{G\}\_\{\\mathrm\{dep\}\}\\left\[\\mathcal\{V\}\\setminus\\left\(\\mathcal\{M\}\_\{\\mathrm\{del\}\}\\cup\\mathcal\{M\}\_\{\\mathrm\{quar\}\}\\right\)\\right\]denote the dependency graph after deleted and quarantined memories are treated as severed nodes\. We define

𝒜ans=\{v∈𝒜∩𝒱step\|v↝vfinal​in​𝒢−\}\.\\mathcal\{A\}\_\{\\mathrm\{ans\}\}=\\left\\\{v\\in\\mathcal\{A\}\\cap\\mathcal\{V\}\_\{\\mathrm\{step\}\}\\;\\middle\|\\;v\\leadsto v\_\{\\mathrm\{final\}\}\\text\{ in \}\\mathcal\{G\}^\{\-\}\\right\\\}\.Thus, an unsupported execution node is answer\-relevant only when its repaired output can still contribute to the regenerated final answer\. Nodes outside𝒜ans\\mathcal\{A\}\_\{\\mathrm\{ans\}\}remain invalidated but are not replayed\.

The executable replay set is obtained by taking the execution closure of the answer\-relevant affected nodes:

𝒱replay=ExecClosure​\(𝒜ans∪\{vfinal\}\)\.\\mathcal\{V\}\_\{\\mathrm\{replay\}\}=\\mathrm\{ExecClosure\}\\left\(\\mathcal\{A\}\_\{\\mathrm\{ans\}\}\\cup\\left\\\{v\_\{\\mathrm\{final\}\}\\right\\\}\\right\)\.The execution closure recursively adds every invalidated executable prerequisite required to recompute a selected node, including affected memory reads, controlling plans, generating tool actions, producer steps, and memory mutation steps\. Safe prerequisites are not replayed and are instead supplied through𝒱preserve\\mathcal\{V\}\_\{\\mathrm\{preserve\}\}\. When a selected node is a tool observation, the execution closure adds its generating tool action rather than the observation itself\. The final answer node is always included so that the repaired execution produces a fresh answer\.

We define the preserved trace context as

𝒱preserve=𝒱step∖\(𝒱inv∪𝒱replay\)\.\\mathcal\{V\}\_\{\\mathrm\{preserve\}\}=\\mathcal\{V\}\_\{\\mathrm\{step\}\}\\setminus\\left\(\\mathcal\{V\}\_\{\\mathrm\{inv\}\}\\cup\\mathcal\{V\}\_\{\\mathrm\{replay\}\}\\right\)\.Therefore,𝒱preserve\\mathcal\{V\}\_\{\\mathrm\{preserve\}\}contains safe prior execution outputs needed as fixed context during replay\. User inputs and active memory records are supplied separately through𝒰\\mathcal\{U\}and the repaired memory store\.

### 3\.6Selective Replay

Given the repair planℛ\\mathcal\{R\}, the selective replay executor first applies the planned memory dispositions\. The active pre\-replay memory store is

ℳ−=ℳ∖\(ℳdel∪ℳquar\)\.\\mathcal\{M\}^\{\-\}=\\mathcal\{M\}\\setminus\\left\(\\mathcal\{M\}\_\{\\mathrm\{del\}\}\\cup\\mathcal\{M\}\_\{\\mathrm\{quar\}\}\\right\)\.Records inℳquar\\mathcal\{M\}\_\{\\mathrm\{quar\}\}may be retained in a separate audit partition, but they are unavailable to retrieval and execution\.

The executor processes𝒱replay\\mathcal\{V\}\_\{\\mathrm\{replay\}\}in the original trace order, thereby preserving the observed execution ordering\. Memory\-lifecycle relations used only for fault resolution are not treated as replay\-precedence constraints\. Nodes in𝒱preserve\\mathcal\{V\}\_\{\\mathrm\{preserve\}\}are supplied as fixed context, while invalidated nodes outside the replay set remain excluded from subsequent execution\.

Each selected step is recomputed using the current repaired memory state and the repaired prefix of the trace\. The memory state is initialized asℳ−\\mathcal\{M\}^\{\-\}and updated as replayed memory mutations are executed\. Invalidated claims are withheld from the replay context until replacements are generated, preventing unsupported outputs from being cited by later steps\. Replayed memory reads retrieve only from the current repaired store, and replayed memory mutations are applied only when they remain justified by the repaired execution\.

Tool use is also replayed selectively\. If the repaired plan still requires tool use, the executor issues a new action and records the resulting observation as a new trace node\. Existing observations are never independently regenerated or treated as reusable when their generating actions have been invalidated\. For side\-effecting tools, re\-execution is performed only through the resettable, idempotent, or compensating interface assumed in the problem setting\.

LetΔ​ℳreplay\\Delta\\mathcal\{M\}\_\{\\mathrm\{replay\}\}denote the ordered sequence of valid memory writes, deletions, updates, and consolidations generated during replay\. The final repaired active memory store is

ℳ′=Apply⁡\(ℳ−,Δ​ℳreplay\),\\mathcal\{M\}^\{\\prime\}=\\operatorname\{Apply\}\\left\(\\mathcal\{M\}^\{\-\},\\Delta\\mathcal\{M\}\_\{\\mathrm\{replay\}\}\\right\),whereApply\\operatorname\{Apply\}executes the replayed memory mutations in trace order and removes records superseded by a valid delete, update or consolidation\.

The final\-answer nodevfinalv\_\{\\mathrm\{final\}\}is always regenerated from the repaired trace\. Let𝒯′\\mathcal\{T\}^\{\\prime\}denote the resulting execution trace andyfinal′y^\{\\prime\}\_\{\\mathrm\{final\}\}the content of the regenerated final answer node\. The executor outputs the corrected answeryfinal′y^\{\\prime\}\_\{\\mathrm\{final\}\}, repaired execution trace𝒯′\\mathcal\{T\}^\{\\prime\}, and repaired active memory storeℳ′\\mathcal\{M\}^\{\\prime\}\.

## 4Controlled Memory\-Repair Benchmark

##### Task Construction

We create 150 tool\-use cases in shopping assistance, travel booking, and customer support, drawing on task structure from common agent benchmarks\(Yaoet al\.[2024](https://arxiv.org/html/2608.10502#bib.bib26); Denget al\.[2023](https://arxiv.org/html/2608.10502#bib.bib27); Xueet al\.[2025](https://arxiv.org/html/2608.10502#bib.bib28)\)\. Each case contains a multi\-turn session, memory store, tools, and paired clean and faulty traces\. At repair time every method receives the faulty store and trace plus diagnosed fault identifiers; the clean trace, expected answer, benign memory labels, and affected\-node annotations are evaluation\-only\.

Following Section[3\.1](https://arxiv.org/html/2608.10502#S3.SS1), at repair time, a method receives a user session𝒰\\mathcal\{U\}, memory storeℳ\\mathcal\{M\}, faulty execution trace𝒯\\mathcal\{T\}, diagnosed faulty memory setℱ\\mathcal\{F\}, and the available tools\. For evaluation, we additionally retain a paired clean execution, an expected repaired answer, benign memory labels, and node\-level annotations of affected state\. The faulty trace captures downstream contamination caused by memory failures, while the clean trace serves only as an evaluation reference and is not exposed to repair methods\. The prompt contracts used to instantiate agent executions are provided in Appendix[B](https://arxiv.org/html/2608.10502#A2)\.

##### Fault Injection

We inject four faulty types motivated by observed memory risks\(Sunilet al\.[2026](https://arxiv.org/html/2608.10502#bib.bib29); Zouet al\.[2025](https://arxiv.org/html/2608.10502#bib.bib11); Hataliset al\.[2023](https://arxiv.org/html/2608.10502#bib.bib30); Chaoet al\.[2026](https://arxiv.org/html/2608.10502#bib.bib7); Chenet al\.[2025](https://arxiv.org/html/2608.10502#bib.bib8); Zhanget al\.[2026](https://arxiv.org/html/2608.10502#bib.bib31); Chenet al\.[2024](https://arxiv.org/html/2608.10502#bib.bib32)\):*poisoned*records contain adversarial or corrupted values;*stale*records survive a failed deletion, update or consolidation;*wrong\-user*records are associated with another user or context; and*summary\-drift*corrupts consolidation output\. A case is retained only if injection changes the final answer and creates at least one downstream effect\. Thus, the benchmark evaluates recovery after diagnosed faults; it does not estimate natural fault prevalence or detector accuracy\. The detailed fault injection strategies are provided in Appendix[D](https://arxiv.org/html/2608.10502#A4)\.

Table[1](https://arxiv.org/html/2608.10502#S4.T1)summarizes the benchmark distribution across domains and primary memory fault types\. Each task is counted once under the fault type defining its benchmark stratum\. Of the 150 tasks, 128 contain a single diagnosed faulty memory, while 22 contain two or three diagnosed faulty memories: 8 in customer support, 7 in shopping, and 7 in travel\.

Table 1:Controlled benchmark statistics by domain and primary memory fault type\.Each task is counted once under the fault type defining its benchmark stratum\. Pois\. and W\-user denote poisoned and wrong\-user faults, respectively, while Drift denotes summary\-drift\.
##### Evaluation Metrics

We measure \(i\)*recovery*, success under the deterministic task oracle; \(ii\)*recurrence*, the fraction of recovered cases in which the same failure reappears when re\-running the last user input; \(iii\) removal of diagnosed faulty memories and preservation of benign memories; \(iv\) claim invalidation F1, computed by micro\-averaging over the gold set of faulty\-trace claim\-step identifiers requiring invalidation or re\-derivation and the set predicted for invalidation or replay; and \(v\) replayed\-step ratio and LLM calls\. Detailed metric definitions are provided in Appendix[E](https://arxiv.org/html/2608.10502#A5)\.

## 5Experiments

Table 2:Main results on the controlled benchmark\.All metrics except LLM count are reported on a 0–1 scale\. Recurrence is computed only over successfully recovered cases; “–” indicates that no case was recovered\. Claim\-inv\. F1 denotes claim invalidation F1\.Table 3:Ablation results on the controlled benchmark\.The rollback planner mainly improves recovery and recurrence, support checking improves selectivity, and selective replay reduces repair cost\.Table 4:Transfer results on the adapted LongMemEval\-V2 subset\.The subset contains 50 externally derived procedural tasks converted into our repair schema\.We evaluate all methods on the controlled benchmark of Section[4](https://arxiv.org/html/2608.10502#S4)and evaluate transfer on an adapted LongMemEval\-V2 procedural subset\. All methods use GPT\-4o and share the same tool environment, task inputs, faulty memory store, failed execution trace, and diagnosed faulty memory identifiers; plan\-based methods are executed by the same rollback executor\. Additional Gemini\-3\.6\-Flash and Qwen3\.6\-27B results are reported in Appendix[F](https://arxiv.org/html/2608.10502#A6)\. All results are from a single run at temperature 0, with seed 42 for GPT\-4o and Qwen3\.6\-27B; Gemini\-3\.6\-Flash does not expose a seed parameter\. Qualitative analyses of representative repair cases in both settings are in Appendix[G](https://arxiv.org/html/2608.10502#A7)\.

We compare against six baselines\. Implementation details are in Appendix[H](https://arxiv.org/html/2608.10502#A8)\.No repairleaves the faulty execution unchanged\. Memory\-centric baselines areFull memory reset,Delete retrieved memories, and aMemAudit\-stylebaseline adapted from post\-hoc memory auditing\(Tanet al\.[2026](https://arxiv.org/html/2608.10502#bib.bib10)\), which uses an oracle\-assisted diagnostic fallback only when its audit candidate set is empty\. Trace\-centric baselines areLLM\-judge repair, inspired by Reflexion\(Shinnet al\.[2023](https://arxiv.org/html/2608.10502#bib.bib9)\)and Self\-Refine\(Madaanet al\.[2023](https://arxiv.org/html/2608.10502#bib.bib33)\), andAgentTrace\-style, adapted from causal graph tracing\(Wang[2026](https://arxiv.org/html/2608.10502#bib.bib34)\)\. Prompts for LLM\-judge repair are provided in Appendix[B](https://arxiv.org/html/2608.10502#A2)\.

### 5\.1Main Results

Table[2](https://arxiv.org/html/2608.10502#S5.T2)shows that our method achieves the highest recovery on the controlled benchmark\. Ours recovers 85\.3% of cases, compared with 77\.3% for LLM\-judge repair and 60\.7% for AgentTrace\-style repair\. This corresponds to absolute gains of 8\.0 and 24\.6 percentage points, respectively\. The gap is larger against memory\-centric baselines\. Ours achieves more than twice the recovery of full memory reset and about 2\.5 times the recovery of Delete retrieved memories and MemAudit\-style repair\. Ours also removes all faulty memories while preserving 100\.0% of benign memories, about 1\.27 times as much benign state as Delete retrieved memories\. These results suggest that dependency\-guided rollback repairs faulty state without broadly deleting useful memory\.

Recurrence gives a more nuanced picture\. Ours reduces recurrence compared with AgentTrace\-style and Delete retrieved memories, but does not obtain the lowest recurrence overall\. Since recurrence is computed only over recovered cases, methods with low recovery can obtain deceptively favorable recurrence values over a smaller subset\. We therefore interpret recurrence together with recovery, faulty memory removal, and preservation\.

Memory\-centric baselines directly edit memory state, but they either delete too much benign memory or fail to remove downstream contamination\. Full memory reset removes faulty memories but destroys all benign memory\. Delete retrieved memories preserves more benign state but recovers far fewer cases than Ours\. And MemAudit\-style repair is conservative but misses many faulty memories and derived effects\. In contrast, trace\-centric methods more accurately identify claim\-level trace state requiring repair\. LLM\-judge repair and AgentTrace\-style obtain higher claim invalidation F1 than Ours on the controlled benchmark\. However, this does not translate into better end\-to\-end repair\. Compared with LLM\-judge repair, Ours achieves higher recovery while reducing replay ratio by 43\.3% and LLM calls by 41\.8%\. AgentTrace\-style is cheaper, but this comes with a 24\.6\-percentage\-point absolute drop in recovery and weaker faulty memory removal\. Overall, Ours offers a favorable recovery–cost trade\-off: compared with LLM\-judge repair, it achieves higher recovery with fewer replayed steps and LLM calls; compared with AgentTrace\-style, it substantially improves recovery and faulty memory removal at the cost of additional replay and LLM calls\.

### 5\.2Ablation Study

Table[3](https://arxiv.org/html/2608.10502#S5.T3)shows that the rollback planner is the component most directly responsible for answer recovery\. Removing it reduces recovery from 85\.3% to 71\.3%, increases recurrence from 26\.6% to 43\.0%, and lowers benign memory preservation from 100\.0% to 99\.6%\. It also raises replay ratio from 12\.3% to 15\.2% and LLM calls from 5\.70 to 9\.19\. Equivalently, the full method uses 38\.0% fewer LLM calls than the variant without the rollback planner\. These results indicate that rule\-guided rollback planning is important not only for identifying what should be invalidated or regenerated, but also for avoiding unnecessary repair operations\.

Support checking primarily improves preservation and selectivity rather than raw recovery\. Removing it increases recovery from 85\.3% to 88\.0% and slightly reduces recurrence from 26\.6% to 26\.5%, but lowers benign memory preservation from 100\.0% to 98\.6%\. It also increases replay ratio by 25\.2% and LLM calls by 14\.2%\. Thus, support checking introduces a small recovery trade\-off in exchange for preserving more independently supported state and reducing unnecessary recomputation\.

Selective replay primarily controls repair cost\. Without it, recovery decreases slightly from 85\.3% to 84\.0%, while recurrence improves from 26\.6% to 7\.1%\. This reduction in recurrence, however, requires substantially broader replay: replay ratio rises from 12\.3% to 75\.5%, and LLM calls increase from 5\.70 to 24\.01\. Relative to this broader\-replay variant, selective replay reduces replay ratio by 83\.7% and LLM calls by 76\.3%\. Selective replay is therefore an efficiency mechanism with an explicit recurrence–cost trade\-off, enabling targeted repair rather than exhaustive recomputation\.

### 5\.3Transfer to Adapted LongMemEval\-V2 Subset

We evaluate an adapted subset of LongMemEval\-V2\(Wuet al\.[2026](https://arxiv.org/html/2608.10502#bib.bib35)\)\. The subset contains 50 externally derived procedural and navigation\-style tasks converted into our repair schema \(Appendix[I](https://arxiv.org/html/2608.10502#A9)\)\. Unlike the controlled benchmark, this subset targets the poisoned faulty type and is predominantly multi\-fault: 5 cases contain one faulty memory, while 45 contain 2–4 faults, as summarized in Table[21](https://arxiv.org/html/2608.10502#A9.T21)\. We therefore treat it as an end\-to\-end transfer stress test on externally derived multi\-fault trajectories rather than a fault\-type coverage study\.

Table[4](https://arxiv.org/html/2608.10502#S5.T4)shows that our method again achieves the highest recovery, 68\.0% against 54\.0% for AgentTrace\-style and 26\.0% for LLM\-judge repair\. In contrast to its controlled benchmark result, our method also obtains the highest claim invalidation F1 on this subset, scoring 0\.669 versus 0\.603 and 0\.369\. These trajectories are dominated by poisoned memories on the answer path, so the claim steps requiring repair largely overlap with the answer\-relevant region selected for replay\. It also ties for the highest faulty memory removal and preserves 99\.3% of benign memories\. AgentTrace\-style keeps a cost advantage and has slightly better recurrence \(3\.7% vs\. 5\.9%\) and preservation \(99\.7% vs\. 99\.3%\), but trails by 14\.0 points in recovery; against LLM\-judge repair, our method is better on every reported transfer metric while using 73\.4% of its replay ratio and 77\.3% of its LLM calls\. That absolute recovery falls from 85\.3% to 68\.0% on externally derived trajectories is the clearest evidence that the controlled results are an upper bound rather than a deployment estimate\. Together with Appendix[F](https://arxiv.org/html/2608.10502#A6), this suggests that our recovery–cost profile is more consistent across settings\.

## 6Conclusion

This paper studies post\-failure memory recovery for memory\-augmented agents\. We propose dependency\-guided rollback repair, which traces contamination through dependency graphs, computes cost\-aware rollback sets, and selectively replays affected steps\. We also build a controlled benchmark spanning shopping, travel, and customer support\. Across both evaluations, our method achieves the highest immediate recovery and a favorable recovery–cost trade\-off\. Remaining controlled\-set gaps in recurrence highlight a key direction for future work: improving recurrence robustness and claim\-state identification without sacrificing selective repair\.

## References

- Dynamic program slicing\.InProceedings of the ACM SIGPLAN 1990 Conference on Programming Language Design and Implementation,pp\. 246–256\.External Links:[Document](https://dx.doi.org/10.1145/93542.93576)Cited by:[§2](https://arxiv.org/html/2608.10502#S2.SS0.SSSx3.p1.1)\.
- H\. Chao, Y\. Bai, R\. Sheng, T\. Li, and Y\. Sun \(2026\)STALE: can llm agents know when their memories are no longer valid?\.arXiv preprint arXiv:2605\.06527\.Cited by:[§1](https://arxiv.org/html/2608.10502#S1.p1.1),[§2](https://arxiv.org/html/2608.10502#S2.SS0.SSSx1.p1.1),[§4](https://arxiv.org/html/2608.10502#S4.SS0.SSS0.Px2.p1.1)\.
- D\. Chen, S\. Niu, K\. Li, P\. Liu, X\. Zheng, B\. Tang, X\. Li, F\. Xiong, and Z\. Li \(2025\)Halumem: evaluating hallucinations in memory systems of agents\.arXiv preprint arXiv:2511\.03506\.Cited by:[§1](https://arxiv.org/html/2608.10502#S1.p1.1),[§2](https://arxiv.org/html/2608.10502#S2.SS0.SSSx1.p1.1),[§4](https://arxiv.org/html/2608.10502#S4.SS0.SSS0.Px2.p1.1)\.
- Z\. Chen, Z\. Xiang, C\. Xiao, D\. Song, and B\. Li \(2024\)Agentpoison: red\-teaming llm agents via poisoning memory or knowledge bases\.Advances in Neural Information Processing Systems37,pp\. 130185–130213\.Cited by:[§2](https://arxiv.org/html/2608.10502#S2.SS0.SSSx1.p1.1),[§4](https://arxiv.org/html/2608.10502#S4.SS0.SSS0.Px2.p1.1)\.
- J\. Cheney, L\. Chiticariu, and W\. Tan \(2009\)Provenance in databases: why, how, and where\.Foundations and Trends in Databases1\(4\),pp\. 379–474\.External Links:[Document](https://dx.doi.org/10.1561/1900000006)Cited by:[§2](https://arxiv.org/html/2608.10502#S2.SS0.SSSx3.p1.1)\.
- X\. Deng, Y\. Gu, B\. Zheng, S\. Chen, S\. Stevens, B\. Wang, H\. Sun, and Y\. Su \(2023\)Mind2web: towards a generalist agent for the web\.Advances in Neural Information Processing Systems36,pp\. 28091–28114\.Cited by:[§4](https://arxiv.org/html/2608.10502#S4.SS0.SSS0.Px1.p1.1)\.
- Y\. Du, B\. Wang, Y\. He, B\. Liang, B\. Wang, Z\. Li, L\. Gui, J\. Z\. Pan, R\. Xu, and K\. Wong \(2026\)MemGuide: intent\-driven memory selection for goal\-oriented multi\-session llm agents\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 30584–30592\.Cited by:[§2](https://arxiv.org/html/2608.10502#S2.SS0.SSSx1.p1.1)\.
- K\. Hatalis, D\. Christou, J\. Myers, S\. Jones, K\. Lambert, A\. Amos\-Binks, Z\. Dannenhauer, and D\. Dannenhauer \(2023\)Memory matters: the need to improve long\-term memory in llm\-agents\.InProceedings of the AAAI Symposium Series,Vol\.2,pp\. 277–280\.Cited by:[§4](https://arxiv.org/html/2608.10502#S4.SS0.SSS0.Px2.p1.1)\.
- A\. Madaan, N\. Tandon, P\. Gupta, S\. Hallinan, L\. Gao, S\. Wiegreffe, U\. Alon, N\. Dziri, S\. Prabhumoye, Y\. Yang,et al\.\(2023\)Self\-refine: iterative refinement with self\-feedback\.Advances in neural information processing systems36,pp\. 46534–46594\.Cited by:[§2](https://arxiv.org/html/2608.10502#S2.SS0.SSSx4.p1.1),[§5](https://arxiv.org/html/2608.10502#S5.p2.1)\.
- C\. Mohan, D\. Haderle, B\. Lindsay, H\. Pirahesh, and P\. Schwarz \(1992\)ARIES: a transaction recovery method supporting fine\-granularity locking and partial rollbacks using write\-ahead logging\.ACM Transactions on Database Systems17\(1\),pp\. 94–162\.External Links:[Document](https://dx.doi.org/10.1145/128765.128770)Cited by:[§2](https://arxiv.org/html/2608.10502#S2.SS0.SSSx3.p1.1)\.
- C\. Packer, V\. Fang, S\. G\. Patil, K\. Lin, S\. Wooders, and J\. E\. Gonzalez \(2023\)MemGPT: towards llms as operating systems\.\.Cited by:[§1](https://arxiv.org/html/2608.10502#S1.p1.1),[§2](https://arxiv.org/html/2608.10502#S2.SS0.SSSx1.p1.1)\.
- J\. S\. Park, J\. O’Brien, C\. J\. Cai, M\. R\. Morris, P\. Liang, and M\. S\. Bernstein \(2023\)Generative agents: interactive simulacra of human behavior\.InProceedings of the 36th annual acm symposium on user interface software and technology,pp\. 1–22\.Cited by:[§1](https://arxiv.org/html/2608.10502#S1.p1.1),[§2](https://arxiv.org/html/2608.10502#S2.SS0.SSSx1.p1.1)\.
- N\. Shinn, F\. Cassano, A\. Gopinath, K\. Narasimhan, and S\. Yao \(2023\)Reflexion: language agents with verbal reinforcement learning\.Advances in neural information processing systems36,pp\. 8634–8652\.Cited by:[§1](https://arxiv.org/html/2608.10502#S1.p3.1),[§2](https://arxiv.org/html/2608.10502#S2.SS0.SSSx4.p1.1),[§5](https://arxiv.org/html/2608.10502#S5.p2.1)\.
- B\. D\. Sunil, I\. Sinha, P\. Maheshwari, S\. Todmal, S\. Mallik, and S\. Mishra \(2026\)Memory poisoning attack and defense on memory based llm\-agents\.arXiv preprint arXiv:2601\.05504\.Cited by:[§4](https://arxiv.org/html/2608.10502#S4.SS0.SSS0.Px2.p1.1)\.
- Z\. Tan, Y\. Yao, H\. Jin, W\. Yu, G\. Wang, M\. Fan, F\. Liu, X\. Zhang, D\. Ma, T\. Yang,et al\.\(2026\)MemAudit: post\-hoc auditing of poisoned agent memory via causal attribution and structural anomaly detection\.arXiv preprint arXiv:2605\.23723\.Cited by:[Appendix H](https://arxiv.org/html/2608.10502#A8.SS0.SSS0.Px5.p1.1),[§1](https://arxiv.org/html/2608.10502#S1.p3.1),[§2](https://arxiv.org/html/2608.10502#S2.SS0.SSSx1.p1.1),[§5](https://arxiv.org/html/2608.10502#S5.p2.1)\.
- G\. Wang, Y\. Xie, Y\. Jiang, A\. Mandlekar, C\. Xiao, Y\. Zhu, L\. Fan, and A\. Anandkumar \(2023\)Voyager: an open\-ended embodied agent with large language models\.arXiv preprint arXiv:2305\.16291\.Cited by:[§1](https://arxiv.org/html/2608.10502#S1.p1.1)\.
- Z\. G\. Wang \(2026\)AgentTrace: causal graph tracing for root cause analysis in deployed multi\-agent systems\.arXiv preprint arXiv:2603\.14688\.Cited by:[Appendix H](https://arxiv.org/html/2608.10502#A8.SS0.SSS0.Px6.p1.1),[§2](https://arxiv.org/html/2608.10502#S2.SS0.SSSx4.p1.1),[§5](https://arxiv.org/html/2608.10502#S5.p2.1)\.
- M\. Weiser \(1984\)Program slicing\.IEEE Transactions on Software EngineeringSE\-10\(4\),pp\. 352–357\.External Links:[Document](https://dx.doi.org/10.1109/TSE.1984.5010248)Cited by:[§2](https://arxiv.org/html/2608.10502#S2.SS0.SSSx3.p1.1)\.
- D\. Wu, Z\. Ji, A\. Kawatkar, B\. Kwan, J\. Gu, N\. Peng, and K\. Chang \(2026\)Longmemeval\-v2: evaluating long\-term agent memory toward experienced colleagues\.arXiv preprint arXiv:2605\.12493\.Cited by:[Appendix I](https://arxiv.org/html/2608.10502#A9.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2608.10502#S1.p6.1),[§5\.3](https://arxiv.org/html/2608.10502#S5.SS3.p1.1)\.
- Z\. Xiong, Y\. Lin, W\. Xie, P\. He, Z\. Liu, J\. Tang, H\. Lakkaraju, and Z\. Xiang \(2026\)How memory management impacts llm agents: an empirical study of experience\-following behavior\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 623–645\.Cited by:[§1](https://arxiv.org/html/2608.10502#S1.p1.1),[§2](https://arxiv.org/html/2608.10502#S2.SS0.SSSx1.p1.1)\.
- W\. Xu, Z\. Liang, K\. Mei, H\. Gao, J\. Tan, and Y\. Zhang \(2026\)A\-mem: agentic memory for llm agents\.Advances in Neural Information Processing Systems38,pp\. 17577–17604\.Cited by:[§2](https://arxiv.org/html/2608.10502#S2.SS0.SSSx1.p1.1)\.
- T\. Xue, W\. Qi, T\. Shi, C\. H\. Song, B\. Gou, D\. Song, H\. Sun, and Y\. Su \(2025\)An illusion of progress? assessing the current state of web agents\.arXiv preprint arXiv:2504\.01382\.Cited by:[§4](https://arxiv.org/html/2608.10502#S4.SS0.SSS0.Px1.p1.1)\.
- S\. Yao, N\. Shinn, P\. Razavi, and K\. Narasimhan \(2024\)τ\\tau\-Bench: a benchmark for tool\-agent\-user interaction in real\-world domains\.arXiv preprint arXiv:2406\.12045\.Cited by:[§4](https://arxiv.org/html/2608.10502#S4.SS0.SSS0.Px1.p1.1)\.
- S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. Narasimhan, and Y\. Cao \(2022\)React: synergizing reasoning and acting in language models\.arXiv preprint arXiv:2210\.03629\.Cited by:[§2](https://arxiv.org/html/2608.10502#S2.SS0.SSSx4.p1.1)\.
- D\. Zhang, Y\. Lin, Z\. Wu, Y\. Sun, B\. Li, D\. Li, and H\. Peng \(2026\)Useful memories become faulty when continuously updated by llms\.arXiv preprint arXiv:2605\.12978\.Cited by:[§4](https://arxiv.org/html/2608.10502#S4.SS0.SSS0.Px2.p1.1)\.
- Z\. Zhang, Q\. Dai, X\. Bo, C\. Ma, R\. Li, X\. Chen, J\. Zhu, Z\. Dong, and J\. Wen \(2025\)A survey on the memory mechanism of large language model\-based agents\.ACM Transactions on Information Systems43\(6\),pp\. 1–47\.Cited by:[§1](https://arxiv.org/html/2608.10502#S1.p1.1)\.
- Y\. Zhao, C\. Dai, M\. Kou, and Y\. Xiu \(2026a\)MEMOREPAIR: barrier\-first cascade repair in agentic memory\.arXiv preprint arXiv:2605\.07242\.Cited by:[§2](https://arxiv.org/html/2608.10502#S2.SS0.SSSx2.p1.1)\.
- Y\. Zhao, B\. Yuan, J\. Huang, H\. Yuan, Z\. Yu, H\. Xu, L\. Hu, A\. Shankarampeta, Z\. Huang, W\. Ni,et al\.\(2026b\)AMA\-bench: evaluating long\-horizon memory for agentic applications\.arXiv preprint arXiv:2602\.22769\.Cited by:[§2](https://arxiv.org/html/2608.10502#S2.SS0.SSSx1.p1.1)\.
- W\. Zhong, L\. Guo, Q\. Gao, H\. Ye, and Y\. Wang \(2024\)Memorybank: enhancing large language models with long\-term memory\.InProceedings of the AAAI conference on artificial intelligence,Vol\.38,pp\. 19724–19731\.Cited by:[§1](https://arxiv.org/html/2608.10502#S1.p1.1)\.
- W\. Zou, R\. Geng, B\. Wang, and J\. Jia \(2025\)\{\\\{poisonedrag\}\\\}: Knowledge corruption attacks to\{\\\{retrieval\-augmented\}\\\}generation of large language models\.In34th USENIX Security Symposium \(USENIX Security 25\),pp\. 3827–3844\.Cited by:[§1](https://arxiv.org/html/2608.10502#S1.p3.1),[§2](https://arxiv.org/html/2608.10502#S2.SS0.SSSx1.p1.1),[§4](https://arxiv.org/html/2608.10502#S4.SS0.SSS0.Px2.p1.1)\.

## Appendix ADependency\-guided Rollback Repair Data Schema

We provide the data schemas used by our implementation\. Each data instance contains a user session record, a memory store, an execution trace, diagnosed faulty memories, a corresponding dependency graph, and a repair plan generated during method execution\. These artifacts correspond to the modules in the main paper: the graph construction module uses lineage fields, cited identifiers, generated memory identifiers, and trace predecessor/successor fields to build typed dependency edges; the rollback planner traverses the graph from diagnosed faulty memories to identify affected claims, downstream memories, and replay candidates; and the plan executor consumes the rollback plan to produce the repaired memory store and repaired trace\.

The clean store is used only as an evaluation reference, the faulty store is the input to the repair method, and the repaired store is the output produced by the plan executor\. These stores share the same memory record schema\. Identifiers are preserved for unchanged memories\.

##### Memory record\.

As shown in Table[5](https://arxiv.org/html/2608.10502#A1.T5), each memory record contains a memory identifier, textual content, status, source metadata, lineage metadata, and an optional trust score\. In our current implementation, the trust score is recorded for analysis only, not used by the repair algorithm\. Future work could estimate it more robustly and integrate it into the support checking module\.

Table 5:Schema of memory records in clean, faulty, and repaired memory stores\. In our implementation,sourcetakes values from user inputs, tool observations, agent summaries, claims, and injected faults, whilefault\_typecovers poisoned, stale, wrong\-user, and summary\-drift memories\.
##### Execution step\.

Table[6](https://arxiv.org/html/2608.10502#A1.T6)summarizes the fields in each execution step, including the step identifier, semantic type, textual content, cited memories or prior steps, tool metadata when applicable, and memory mutation metadata when applicable\.

Table 6:Schema of execution steps in clean, faulty, and repaired traces\.step\_typecovers memory reads, claims, planning steps, tool actions and observations, final answers, and memory write/delete/update/consolidation operations\.
##### Dependency graph\.

Each graph node has a node identifier, node kind, semantic type, and metadata, and each graph edge has a source node, target node, and one or more typed labels\. Detailed fields are presented in Tables[7](https://arxiv.org/html/2608.10502#A1.T7)and[8](https://arxiv.org/html/2608.10502#A1.T8), with edge labels listed in Table[9](https://arxiv.org/html/2608.10502#A1.T9)\.

The graph is constructed from the memory store and execution trace, so memory nodes correspond to memory records, step nodes correspond to execution step records, and edge endpoints resolve to nodes in the same task instance\. Chain edges are consistent withprev\_step\_idsandnext\_step\_ids; citation and support edges are consistent withused\_ids, while production edges are consistent withgenerated\_memory\_ids\. Edge directions follow provenance and execution flow\. Supporting or antecedent nodes point to dependent nodes, and earlier execution steps point to later steps\.

Table 7:Schema of dependency graph nodes\.Table 8:Logical schema of dependency graph edges\. In the implementation, edges are stored as adjacency maps from source nodes to target nodes with a list of typed labels\.Table 9:Dependency graph edge labels, including their direction and meaning\.
##### Repair plan\.

Repair plans contain task identifiers and detailed instructions for the plan executor in Table[10](https://arxiv.org/html/2608.10502#A1.T10)\. Memories are deleted or quarantined, while selected steps are replayed\. Invalidated claims and redundant steps support state reconstruction, and suspicious steps are retained for audit rather than replayed by the default executor\.

Plan fields are interpreted as disjoint action sets unless otherwise specified: a memory cannot appear in bothdelete\_memory\_idsandquarantine\_memory\_ids, and a trace step cannot simultaneously appear inreplay\_step\_ids,preserve\_step\_ids, andredundant\_step\_ids\. All identifiers in the plan must resolve to memories or trace steps in the same task instance\.

Table 10:Schema of rollback plan artifacts\.

## Appendix BPrompt Contracts and Templates

Table 11:Domain instructions used to instantiate agent prompts\.Table 12:Prompt contracts for the base memory\-augmented agent\.Table 13:Prompt contracts for rollback replay and the LLM\-judge repair baseline\.We summarize the structured prompt contracts used by the base agent, our rollback executor, and the LLM\-judge repair baseline\. Because these components are implemented through structured prompting, the contracts are part of the experimental protocol\. They define what information each component can access, what evidence it must cite, and what JSON outputs it must produce\. Runtime prompts instantiate these reusable contracts with domain instructions, user turns, retrieved memories, prior turns, tool schemas, tool observations, repaired state artifacts, and required output schemas\. We report the contracts rather than fully instantiated prompts\.

##### Domain Prompts\.

The executor prepends one domain instruction to the claim, planning, response, and memory mutation prompts\. Table[11](https://arxiv.org/html/2608.10502#A2.T11)lists the controlled benchmark instructions and the separate instruction used for the adapted LongMemEval\-V2 subset\. These domain prompts are intentionally concise so that domain framing does not encode repair specific behavior\.

##### Agent Prompts\.

The base memory\-augmented agent decomposes each turn into claim generation, planning and tool use, response generation, and post\-execution memory mutation\. Each stage returns strict JSON with provenance identifiers, which are recorded in the trace and later used by graph construction and repair\. The planning output is parsed by the executor into concrete tool action steps\. Corresponding tool results are recorded as tool observation steps and made available to response generation and memory mutation\.

##### Rollback Repair and Baseline Prompts\.

Replay prompts are selected by the rollback plan and instantiated over the repaired active state\. They return strict JSON with evidence ids\. For targeted replay, old actions or memory mutations are shown only as untrusted hints, and the model must produce replacements grounded in repaired evidence\.

## Appendix CRule\-Guided Rollback Decisions

Table 14:Rule\-guided rollback decisions\.Memory dispositions determine which records remain active, whereas trace decisions determine which prior outputs are preserved, invalidated, or replayed\.Table[14](https://arxiv.org/html/2608.10502#A3.T14)summarizes the principal repair rules applied by the rollback planner\. Memory dispositions determine which records remain active, while trace actions determine which prior outputs are invalidated, replayed, or preserved\.

## Appendix DFault Injection Details

We inject faulty memories into clean benchmark instances to construct post\-failure memory states\. Fault selection is manifest\-driven: for each task, a manifest specifies one or more fault specifications, including the fault type, target memory or target mutation step when applicable, and optional corrupted content\. The injector does not score candidate memories or run the agent\. Instead, it rewrites the clean memory store and cut trace so that the replay starts from a seeded faulty state\. When multiple faults are specified for the same task, they are applied sequentially and merged into a single faulty instance\.

Each manifest row contains a task identifier and an ordered list of fault specifications:

r=\(τ,\[f1,…,fk\]\)\.r=\(\\tau,\[f\_\{1\},\\ldots,f\_\{k\}\]\)\.Here,τ\\taudenotes the task identifier, and faults are applied sequentially in the listed order\. Each fault specificationfif\_\{i\}records the fault type and any fault specific fields needed for injection, such as the target memory, target mutation step, corrupted content, or wrong\-user content\. Corrupted content is used for manifest specified corruptions, while wrong\-user content is used only for wrong\-user faults\.

##### Poisoned\.

For poisoned memory faults, the manifest names a target memory produced by a memory write step\. The injector replaces the clean memory content with manifest provided corrupted content when available; otherwise, it corrupts a concrete value using deterministic domain vocabularies or numeric shifts\. The faulty record receives a new faulty identifier, is marked active, and keeps the original temporal position and source metadata while being stamped with injected fault provenance\. The producing trace step is also rewritten so that its generated memory id and textual content are consistent with the corrupted memory\.

For example, in a shopping task, a clean memory such as “The user prefers black running shoes” may be corrupted into “The user prefers red running shoes\.” The corresponding memory write step is rewritten to generate the faulty memory id and the corrupted content, so the seeded trace and memory store remain consistent\.

##### Stale\.

To model stale memory faults, the injector simulates a failed memory mutation in which an old memory that should have been deleted or superseded remains active\. It identifies the delete, update, or consolidation step that consumed the target memory specified in the manifest\. The consumed memory inputs are reactivated as stale faulty records, while the new memories that would have been produced by the mutation, together with their downstream lineage when applicable, are removed from the seeded store\. The consuming step is kept in the cut trace but rewritten so that generated memories are removed and references point to the surviving stale records\.

A clean trace may update a memory stating “The user prefers economy flights” into a new memory stating “The user prefers business class flights\.” A stale injection reactivates the old economy flight preference and removes the new business class preference from the seeded store\. During replay, this causes the agent to behave as if the user still prefers economy flights\.

##### Wrong\-user\.

Wrong\-user faults are injected by fabricating a new active memory from the manifest provided wrong\-user content\. This memory has no producing trace step and is treated as an injected root memory store fault\. It is inserted at the beginning of the memory timeline, and replay starts from the first turn so that the wrong\-user memory can influence subsequent execution\.

For example, a travel task for the current user may receive an injected memory stating “The user prefers hotels with airport shuttle service,” even though this preference belongs to a different user\. Since the injected memory has no producing step in the current trace, it is treated as a root memory store fault and can affect replay from the first turn\.

##### Summary\-drift\.

Summary\-drift faults use the same injection mechanism as poisoned memories, but they specifically target derived memories produced by summarization or consolidation\. This models an inaccurate summary or consolidation output whose content differs from the clean memory while appearing as a confident derived memory\. When specified, additional derived memories can be dropped so that the seeded store reflects the drifted summary lineage\.

In a customer support scenario, two clean memories may state that “the user reported defective headphones” and “the user asked about refund eligibility after returning the item\.” A correct consolidation should summarize them as “The user returned defective headphones and is seeking a refund,” but summary\-drift changes the derived memory to “The user wants to keep the headphones and receive a replacement\.” This drifted summary may cause the agent to create the wrong support ticket or choose an incorrect resolution policy\.

## Appendix EMetrics Details

We evaluate each repair method at the case level and report ratio\-based metrics as percentages unless otherwise stated\. Memory, claim, and replay metrics use micro\-aggregation: their numerators and denominators are summed across cases before the final ratio is computed, thereby avoiding equal weighting of cases containing different numbers of memories, claims, or trace steps\.

##### Recovery\.

Recovery measures whether the repaired agent produces a correct final outcome according to a deterministic task oracle\. For the controlled shopping, travel, and customer support tasks, a case succeeds if the final answer contains all required facts\. For adapted tasks, we apply the answer evaluator from the original LongMemEval\-V2 benchmark\. LetNsuccN\_\{\\mathrm\{succ\}\}andNfailN\_\{\\mathrm\{fail\}\}be the numbers of cases labeled as success and failure, respectively\. We compute

Recovery=NsuccNsucc\+Nfail\.\\mathrm\{Recovery\}=\\frac\{N\_\{\\mathrm\{succ\}\}\}\{N\_\{\\mathrm\{succ\}\}\+N\_\{\\mathrm\{fail\}\}\}\.

##### Recurrence\.

Recurrence measures whether a repaired state remains correct when it is reused and is evaluated only on cases that pass the recovery oracle\. For each successfully recovered case, we instantiate a fresh agent runtime, clear its existing state, load only the active memories from the repaired memory store, and rerun the final user input\. The outcome is labeled as recurrence if its answer fails the deterministic oracle\. LetNrecN\_\{\\mathrm\{rec\}\}andNno​\-​recN\_\{\\mathrm\{no\\text\{\-\}rec\}\}denote the numbers of decided recurrence and no\-recurrence probes, respectively\. We report

Recurrence=NrecNrec\+Nno​\-​rec\.\\mathrm\{Recurrence\}=\\frac\{N\_\{\\mathrm\{rec\}\}\}\{N\_\{\\mathrm\{rec\}\}\+N\_\{\\mathrm\{no\\text\{\-\}rec\}\}\}\.Because recurrence is conditioned on successful recovery, the two metrics should be interpreted jointly\.

##### Faulty removal\.

Faulty removal measures the recall of source faulty memories removed by the repair method\. For casecc, letFcF\_\{c\}be the set of source faulty memory ids specified by the ground truth, and letDcD\_\{c\}contain memories deactivated by the repair or selected for deletion or quarantine:

Dc=Dcdeactivated∪Dcdelete∪Dcquarantine\.D\_\{c\}=D\_\{c\}^\{\\mathrm\{deactivated\}\}\\cup D\_\{c\}^\{\\mathrm\{delete\}\}\\cup D\_\{c\}^\{\\mathrm\{quarantine\}\}\.We compute the micro\-aggregated score

FaultyRemoval=∑c\|Fc∩Dc\|∑c\|Fc\|\.\\mathrm\{FaultyRemoval\}=\\frac\{\\sum\_\{c\}\|F\_\{c\}\\cap D\_\{c\}\|\}\{\\sum\_\{c\}\|F\_\{c\}\|\}\.Thus, the metric credits repair\-attributable removal decisions rather than merely checking whether a memory is absent from the final active store\. AgentTrace\-style achieves a Faulty Removal score of 0 because it focuses exclusively on trace repair and leaves the memory store unchanged\.

##### Benign preservation\.

Benign preservation measures how much non\-faulty memory state survives the repair\. For casecc, letBcB\_\{c\}denote the ground\-truth benign memory ids\. A benign memory that was active when repair began is considered preserved if either its original id or a successor reachable through the memory id remapping chain remains active after repair\. A benign memory that was already inactive is considered preserved as long as the repair does not target it for deactivation, deletion, or quarantine\. IfPc⊆BcP\_\{c\}\\subseteq B\_\{c\}is the set of preserved benign memories, we compute

BenignPreservation=∑c\|Pc\|∑c\|Bc\|\.\\mathrm\{BenignPreservation\}=\\frac\{\\sum\_\{c\}\|P\_\{c\}\|\}\{\\sum\_\{c\}\|B\_\{c\}\|\}\.

##### Claim\-invalidation F1 \(↑\\uparrow\)\.

Claim\-invalidation F1 evaluates whether the repair invalidates exactly the trace claims affected by the fault\. For casecc, letGcG\_\{c\}be the ground\-truth set of affected trace steps whose type isclaim, and letIcI\_\{c\}be the set of claim ids actually invalidated by the repair\. We first micro\-aggregate

TP=∑c\|Gc∩Ic\|,FP=∑c\|Ic∖Gc\|,FN=∑c\|Gc∖Ic\|,\\mathrm\{TP\}=\\sum\_\{c\}\|G\_\{c\}\\cap I\_\{c\}\|,\\qquad\\mathrm\{FP\}=\\sum\_\{c\}\|I\_\{c\}\\setminus G\_\{c\}\|,\\qquad\\mathrm\{FN\}=\\sum\_\{c\}\|G\_\{c\}\\setminus I\_\{c\}\|,and then report

ClaimInvF1=2​T​P2​T​P\+FP\+FN\.\\mathrm\{ClaimInvF1\}=\\frac\{2\\mathrm\{TP\}\}\{2\\mathrm\{TP\}\+\\mathrm\{FP\}\+\\mathrm\{FN\}\}\.The score for memory\-centric baselines is always 0 because they only focus on memory store repair and no claim steps are invalidated at all\.

##### Replay ratio\.

Replay ratio measures the fraction of the original trace that is actually regenerated\. LetScS\_\{c\}be the total number of steps in the original trace andRcR\_\{c\}the number of original target steps that are both selected for replay and successfully mapped to generated replacements\. We report

ReplayRatio=∑cRc∑cSc\.\\mathrm\{ReplayRatio\}=\\frac\{\\sum\_\{c\}R\_\{c\}\}\{\\sum\_\{c\}S\_\{c\}\}\.Auxiliary replacement steps generated as a consequence of replay, such as tool observations, do not increase the numerator unless they are themselves original replay targets\. The metric therefore measures the selectively replayed portion of the original execution rather than the total length of the repaired trace\. Although memory\-centric methods modify only the memory store, they replay the final user turn to regenerate the response under the repaired memory state, resulting in a nonzero replay ratio\. MemAudit\-style has a lower ratio because it emits a no\-op and skips final turn replay when no candidate memory is found\.

##### LLM count\.

LLM count reports the average number of LLM calls made by the repair procedure per case\. For our method, the per\-case count is

Lc=Lcdecision\+Lcreplay,L\_\{c\}=L\_\{c\}^\{\\mathrm\{decision\}\}\+L\_\{c\}^\{\\mathrm\{replay\}\},where the two terms count calls used to make repair decisions and calls used to regenerate replayed steps, respectively\. ForNNevaluated cases, we report

LLMCount=1N​∑c=1NLc\.\\mathrm\{LLMCount\}=\\frac\{1\}\{N\}\\sum\_\{c=1\}^\{N\}L\_\{c\}\.This count excludes calls made during the original faulty execution, deterministic evaluation, and subsequent recurrence probing\.

## Appendix FAdditional Results

Table 15:Per\-domain results for our method\.The method remains effective across all three domains\. Shopping has the lowest recovery\.Table 16:Per\-fault\-type results for our method\.Fault types are grouped by primary fault type\. Summary\-drift and stale faults have higher recovery, while wrong\-user faults are harder to recover and require more LLM calls\.Table 17:Backbone sensitivity on the controlled benchmark\.Supplemental Gemini\-3\.6\-Flash and Qwen3\.6\-27B runs evaluate Ours together with the closest LLM\-based and trace\-based repair competitors\.This appendix reports additional controlled benchmark results\. We focus on slices that explain where recovery remains difficult, how repair behavior changes across fault types, how results vary across LLM backbones, and how effectiveness relates to replay and LLM cost\.

##### Per\-domain results\.

Table[15](https://arxiv.org/html/2608.10502#A6.T15)shows that performance varies across domains\. Customer support achieves the highest recovery at 94\.3%, while shopping is the hardest domain, with recovery dropping to 72\.9%\. Shopping also has the highest recurrence at 40\.0%, suggesting that repaired answers are more likely to remain vulnerable to downstream state errors in this domain\. We attribute this difficulty to fine\-grained product distinctions, including product ids, sellers, prices, inventory, and user preferences, where a small reconstruction error can change the selected item or seller\.

##### Per\-fault\-type results\.

Table[16](https://arxiv.org/html/2608.10502#A6.T16)groups tasks by primary fault type\. Summary\-drift has the highest recovery at 95\.0% and the highest claim invalidation F1 at 0\.870, likely because drifted summaries often have compact provenance through update or consolidation steps\. Wrong\-user faults are the hardest, with recovery dropping to 70\.6% and the highest LLM count, reflecting that wrong\-user memories can look like coherent personalized records unless the trace exposes a conflict with the current user’s history\. Stale faults show a different pattern\. Recovery is high at 94\.1%, but recurrence is also high at 78\.1%, indicating that final answer repair can succeed while stale state suppression remains fragile\. Poisoned faults have the lowest claim invalidation F1, but recurrence is 0 among recovered cases, suggesting that once the corrupted value is removed or refreshed, the repaired state is less likely to reintroduce the same failure\.

##### Backbone sensitivity\.

Table[17](https://arxiv.org/html/2608.10502#A6.T17)reports supplemental controlled benchmark runs with Gemini\-3\.6\-Flash and Qwen3\.6\-27B\. We run all three LLMs with temperature 0 and seed 42, except that Gemini\-3\.6\-Flash does not accept a seed parameter\. The goal is to test whether the main findings depend on a particular LLM backbone, rather than to run the full comparison across all repair paradigms\. We therefore compare Ours with LLM\-judge repair, the closest LLM\-based competitor, and AgentTrace\-style, the closest trace\-based competitor\. Compared with GPT\-4o, recovery decreases for all three methods under the additional backbones, with the largest drops under Qwen3\.6\-27B\. Under Gemini, LLM\-judge repair slightly exceeds Ours in recovery, 0\.740 versus 0\.733, corresponding to only one additional recovered case\. However, Ours has lower recurrence, 0\.091 versus 0\.108, perfect benign preservation, and lower replay and LLM cost\. The lower benign preservation of LLM\-judge repair under Gemini and Qwen reflects occasional over\-repair, where the model removes or quarantines non\-target memories\. Overall, these supplemental results support the main experimental conclusions: backbone choice affects absolute recovery and repair\-plan quality, but dependency\-guided rollback continues to provide a stronger balance of recovery, recurrence, preservation, and cost\.

##### Repair cost breakdown\.

Table[18](https://arxiv.org/html/2608.10502#A6.T18)shows that our method is not simply buying recovery with more computation\. Compared with LLM\-judge repair, our method improves recovery from 77\.3% to 85\.3% while using less than half the total tokens and fewer LLM calls\. AgentTrace\-style is cheaper, with a lower replay ratio of 9\.1%, but its recovery drops to 60\.7%, showing the cost of under\-repairing persistent memory contamination\. The ablations clarify the main cost driver\. Removing selective replay raises the replay ratio from 12\.3% to 75\.5%, while recovery does not improve over the full method\. Removing support checking mainly affects selectivity, and removing the rollback planner hurts both recovery and efficiency\.

Table 18:Repair cost breakdown\.Our method achieves higher recovery than LLM\-judge repair while using fewer tokens and fewer LLM calls\. The ablations show that selective replay is the main mechanism preventing repair cost from increasing sharply\.
##### Sensitivity to fault and memory scale\.

Table[19](https://arxiv.org/html/2608.10502#A6.T19)reports robustness slices by faulty memory count and benign active memory store size\. These are not controlled scaling laws because bucket sizes and fault compositions vary, but the method does not collapse on the available multi\-fault cases\. Recovery is 85\.2% for one\-fault cases and 82\.4% for two\-fault cases, while the three\-fault bucket remains recoverable but contains only five cases\. Memory store size is also non\-monotonic\. Recovery drops to 77\.8% in the 6–8 bucket but rises to 91\.4% in the 9–11 bucket\. Replay ratio remains modest across buckets, ranging from 10\.8% to 15\.5%\. These trends suggest that repair difficulty depends more on dependency shape, fault semantics, and whether the faulty memory lies on a load\-bearing answer path than on memory store size alone\.

Table 19:Robustness slices by faulty memory count and memory store size\.Memory store size is measured by the number of benign active memories\. These slices are intended as robustness checks rather than scaling claims, since bucket sizes and fault compositions are not controlled independently\.

## Appendix GCase Study

We present one running example, two compact cases from the controlled benchmark, and one multi\-fault case from the adapted LongMemEval\-V2 subset\. The customer support example illustrates the full repair pipeline, while the travel and shopping cases highlight two complementary mechanisms: cleaning a wrong\-user root memory together with contaminated derived memories, and replaying a load\-bearing tool\-call path after a poisoned factual memory\. The final case illustrates how multiple poisoned memories can reinforce the same error across different levels of abstraction, requiring both persistent memory cleanup and trace\-level replay for complete repair\.

### G\.1Running Example: Customer Support Summary\-Drift Case

We usecustomer\_support\_task\_036because the fault propagates beyond the final answer into downstream claims and memory mutations\.

##### Task context\.

> User input:Using only the latest concise support case summary you stored, tell me the order status and refund window together\. Faulty answer:The order status forORD00169is shipped, and the refund window is45 days\. Expected repaired answer:OrderORD00169is shipped, and the Home refund window is30 days\.

##### Fault provenance\.

The failure is caused by a drifted memory update\. The prior summarym\_019correctly records that orderORD00169is shipped and has a Home refund window of30 days\. The faulty updated summarym\_f026changes the refund window to45 days\. Given diagnosed faulty memorym\_f026, the provenance module follows the incomingproduceedge to update steps\_132, which consumedm\_019, invalidated it, and produced the drifted summary\. This identifies the fault as summary\-drift during memory update rather than an isolated root memory store fault\.

##### Affected state and support checking\.

Forward tracing froms\_132andm\_f026identifies 21 affected nodes: 5 memory nodes, 4 memory read steps, 4 claim steps, 4 final answer steps, 3 memory write steps, and 1 memory update step\. The answer\-relevant region contains the faulty final answers\_148, which depends on a downstream claim using the drifted refund window\.

Support checking prevents all reachable nodes from being treated as unsafe\. After removing the faulty summary, the checker markss\_132as independently supported, while downstream affected claims and answer steps remain unsupported\. It also separates affected but supported derived memories from unsafe ones, preservingm\_027,m\_028,m\_029, andm\_030because they remain independently supported despite being reachable from the faulty summary\.

##### Rollback and replay\.

The planner deletesm\_f026, replays the faulty summary producer and the answer\-relevant downstream computation, and leaves unsupported but not load\-bearing steps suspicious rather than replaying them\. Support verdicts are therefore not hard constraints\.s\_132is replayed because it produced the faulty summary, even though its source evidence can be independently supported\. Conversely,s\_137is left suspicious because it is unsupported but not load\-bearing for the corrected final answer\.

After replay, the faulty summary is replaced by corrected memorym\_n034, which states that the Home refund window is30 days\. Unrelated support memories are preserved,m\_028remains active, and affected side memories are either preserved or regenerated according to their support status\.

##### Corrected output and baseline behavior\.

The repaired final answer states that orderORD00169is shipped and that the Home refund window is30 days\. No repair propagates the drifted summary\. MemAudit\-style preserves the drifted summary because memory auditing alone does not force replay of the affected summary and answer path\. Delete retrieved memories removes useful context and yields an incomplete answer, while full reset deletes the support history needed to answer from memory\. LLM\-judge repair and AgentTrace\-style recover the answer, but do not produce the same selective rollback account over the faulty producer and downstream state\.

### G\.2Derived Faulty Memories: Travel Wrong\-User Memory Case

Instancetravel\_task\_048asks for the airline of the exact confirmed active flight\. The faulty answer is Qantas, while the expected repaired answer is ANA\. The source fault is a wrong\-user memorym\_f000stating that the active flight airline was corrected to Qantas rather than ANA\. Since this memory has no producing step in the current trace, our method treats it as a root memory store fault\.

The wrong\-user memory also contaminates a downstream active flight summary\. In the clean store,m\_008records flightF0028with airline ANA and price$275\.99\. In the faulty run, the derived summary is rewritten with Qantas\. Thus the final answer is supported both by the wrong\-user root memory and by a contaminated derived memory\. Our method deletesm\_f000, remaps the contaminated active flight summary to corrected memorym\_n012, and recovers ANA while preserving unrelated travel memories\.

This case illustrates why persistent memory state repair is necessary\. No repair and AgentTrace\-style leave the root wrong\-user memory and derived summary active, MemAudit\-style removes the derived summary but misses the root memory, and LLM\-judge repair fixes the immediate answer but misses the contaminated derived memory, leading to recurrence\. Full reset and delete retrieved memories can recover the answer, but only by sacrificing substantially more benign state\. Delete retrieved memories preserves only 6 in 10 benign memories, while full reset deletes all 12 active memories\.

### G\.3Tool Call Replay: Shopping Poisoned Memory Case

Instanceshopping\_task\_014asks the agent to compare current prices for the remembered product\. The faulty answer selects StoreB at$66\.95and the expected repaired answer selects StoreC at$63\.99\. The source faultm\_f003poisons the StoreC price for productP0199, raising it to$88\.67\. This makes the faulty answer locally consistent with corrupted price evidence, since StoreB appears cheaper than StoreC\.

Reliable repair therefore requires more than deleting the poisoned memory or editing the final response\. The load\-bearing comparison path must be replayed so thatcompare\_price\(P0199\)refreshes the tool\-derived evidence\. Our method deletesm\_f003, replays the answer\-relevant path including the comparison tool call, regenerates the price comparison memory, and answers StoreC at$63\.99\. The replayed tool result reports StoreC at$63\.99, StoreA at$67\.00, and StoreB at$66\.95, which flips the final comparison while preserving other active product memories\.

All baselines fail this case\. Memory\-centric methods either leave the poisoned price active or remove too much product context\. LLM\-judge repair deletes the poisoned memory, and AgentTrace\-style avoids keeping it active, but neither recovers the corrected tool\-derived evidence needed to change the final comparison\. The case shows why dependency\-guided rollback must sometimes replay the tool action that produced or refreshed contaminated evidence\.

### G\.4Multi\-Fault: Adapted LongMemEval\-V2 Case

We include one adapted LongMemEval\-V2 case to illustrate how multiple poisoned memories can jointly contaminate execution\. Instancelme\_task\_025\_901b7d17asks the agent to identify the new header action button that appears after opening a ServiceNow knowledge article from the search\-results page\. The expected answer isAttach to Private Task, while the faulty execution answersAdd to Private Task\.

This case contains three poisoned memories that encode the same wrong conclusion at different abstraction levels\. Memorym\_f005corrupts the retrieved evidence state by recording the article\-page header button asAdd to Private Task\. Memorym\_f006turns this corrupted observation into a derived factual statement, claiming thatAdd to Private Taskis the label visible near the article heading and absent from the search\-results context\. Memorym\_f009further consolidates the error into a compact answer memory\. The new article\-page header action label isAdd to Private Task\. Thus the final answer is not caused by a single isolated poisoned fact\. The same wrong label is reinforced across raw evidence, extracted observation, and compact answer memory\.

Starting from the diagnosed faulty memories, our method deletesm\_f005,m\_f006, andm\_f009\. During rollback, it invalidates the affected claim stepss\_046ands\_049, replays the answer\-relevant path, and reruns the evidence lookup needed to refresh the article\-page observation\. The repaired execution generates replacement memoriesm\_n010,m\_n011, andm\_n012\. These replacements consistently record the corrected labelAttach to Private Task: the repaired raw evidence memory stores the corrected article\-page button, the repaired derived observation states that this label appears on the article page, and the repaired compact answer memory identifies it as the new header action\. The corrected final answer is thereforeAttach to Private Task\.

This case highlights why multi\-level poisoned memories require both memory\-state cleanup and trace\-level replay\. No repair keeps the poisoned memories active and repeats the faulty answer\. Memory\-centric baselines delete the poisoned memories but fail to recover the answer\. MemAudit\-style answersSubscribe, while full reset and delete retrieved memories answerEditafter losing useful context\. LLM\-judge repair also deletes all three poisoned memories but still regeneratesAdd to Private Task\. AgentTrace\-style recovers the final answer through trace\-level replay, but without the same explicit cleanup account over all jointly poisoned memory levels\. Our method recovers the answer while deleting all three poisoned memories, preserving all benign memories, and regenerating consistent replacement memories for the raw evidence, derived observation, and compact answer\.

## Appendix HBaseline Implementation Details

We summarize the baseline implementations used in our evaluation\. Since prior methods do not directly target post\-failure repair for memory\-augmented executions, we adapt each method to the same repair interface\. Each method receives the failed trace, faulty memory store, diagnosed faulty memory ids, and final turn task input, then outputs either a memory edit decision or a rollback plan executed by the shared replay executor\. Table[20](https://arxiv.org/html/2608.10502#A8.T20)summarizes the configurations\.

Table 20:Comparison of repair baselines\.Baselines differ in whether they repair memory state, roll back faulty trace state, or use dependency structure\.##### No Repair\.

This lower\-bound condition keeps the failed run unchanged, including the faulty memory store, faulty trace, and original final answer\. It emits an empty repair plan and performs no cleanup, deletion, rollback, or replay\.

##### Full Memory Reset\.

This memory\-centric baseline removes all active memories before replaying the final user turn under the same agent lifecycle\. The trace prefix before the final turn is kept fixed, and the regenerated final turn trace is normalized to the same artifact schema as other methods\. This strategy removes all potentially contaminated memory, but also deletes benign personalization state\.

##### Delete Retrieved Memories\.

This baseline deletes active memories that appear in the affected region of the failed execution before the final turn\. Starting from the affected subgraph seeded by the diagnosed faulty memories, it collects active memories retrieved or referenced by relevant non\-retrieval steps, where non\-retrieval steps refer to all steps except memory read steps, then deletes those memories and replays the final user turn\. It does not perform independent\-support checking, derived memory cleanup, or answer\-relevant rollback optimization\.

##### LLM\-Judge Repair\.

This baseline asks an LLM to predict a rollback plan from a compact view of the failed execution, including session turns, diagnosed faulty memory ids, compact memory records, compact trace steps, the wrong final answer, valid id sets, and the executor output schema\. The output is constrained to a JSON rollback plan with fields for memory deletion, memory quarantine, replay steps, preserved steps, and suspicious steps\. The plan must reference only provided memory and step identifiers and is required to delete every diagnosed faulty memory\.

We validate each generated plan before execution\. Plans are rejected if they contain invalid JSON, malformed fields, invalid ids, overlapping step dispositions, or missing deletion of diagnosed faulty memories\. Validation failures are retried with the error appended to the prompt\. If no valid plan is produced within the retry budget, the case is marked as failed\. Valid plans are executed by the same rollback executor as our method, isolating the effect of the planning strategy\.

##### MemAudit\-style\.

We adapt MemAudit\-style post\-hoc memory auditing\(Tanet al\.[2026](https://arxiv.org/html/2608.10502#bib.bib10)\)as a memory deletion selector\. The candidate pool contains active memories before the final user turn that are used by downstream non\-retrieval steps\. If this pool is empty, we fall back to diagnosed source fault memories that remain active\. We report this fallback as an oracle\-assisted diagnostic fallback used only to avoid an empty audit candidate set, not as an unsupervised MemAudit signal\. Candidate memories are scored using a memory consistency anomaly signal and a counterfactual influence signal\. The consistency graph compares memories with matching entity and fact fields\. Matching values induce support edges, while different values induce contradiction edges\.

For candidate memorymm, the anomaly score is

anomaly​\(m\)=contradiction​\(m\)support​\(m\)\+contradiction​\(m\),\\mathrm\{anomaly\}\(m\)=\\frac\{\\mathrm\{contradiction\}\(m\)\}\{\\mathrm\{support\}\(m\)\+\\mathrm\{contradiction\}\(m\)\},with score0when the denominator is zero\. In the oracle\-assisted diagnostic variant, if benchmark provided faulty content directly matches a memory, that memory is assigned anomaly score1\.01\.0\. Counterfactual influence is computed by deletingmm, rerunning the final user turn, and scoring the regenerated trace with the deterministic task oracle:

influence​\(m\)=max⁡\(0,S−m−Sorig\)1−Sorig,\\mathrm\{influence\}\(m\)=\\frac\{\\max\(0,S\_\{\-m\}\-S\_\{\\mathrm\{orig\}\}\)\}\{1\-S\_\{\\mathrm\{orig\}\}\},whereSorigS\_\{\\mathrm\{orig\}\}is the original faulty trace score andS−mS\_\{\-m\}is the score after deletingmm\. If the original score is already11, or if the rerun fails, the influence score is set to0\. The final score is

score​\(m\)=α⋅influence​\(m\)\+\(1−α\)⋅anomaly​\(m\),\\mathrm\{score\}\(m\)=\\alpha\\cdot\\mathrm\{influence\}\(m\)\+\(1\-\\alpha\)\\cdot\\mathrm\{anomaly\}\(m\),withα=0\.5\\alpha=0\.5\. Memories with score above0\.50\.5are deleted together, and only the final user turn is replayed\. This baseline does not roll back trace nodes, invalidate claims, repair derived memory lineage, or select answer relevant replay frontiers\.

##### AgentTrace\-style\.

We adapt AgentTrace\-style causal trace localization\(Wang[2026](https://arxiv.org/html/2608.10502#bib.bib34)\)as a trace\-centric repair baseline\. The method receives the same diagnosed faulty memory ids as other methods, constructs faulty memory provenance and affected regions, scores candidate trace steps, selects the highest ranked step as the root cause, and converts it into a replay plan executed by the shared rollback executor\. We use this adaptation only to choose replay regions, not to implement a full persistent memory repair workflow\.

Candidate steps are restricted to memory read, claim, plan, tool action, and final answer steps that are backward reachable from the erroneous final answer and fault\-relevant\. A step is fault\-relevant if it lies in the affected trace region, directly uses faulty memory, belongs to the fault dependent suffix, or is the final answer step\. Each candidatevvreceives

r​\(v\)=\\displaystyle r\(v\)=0\.40​fback​\(v\)\+0\.25​faff​\(v\)\+0\.20​ffault​\(v\)\\displaystyle 40f\_\{\\mathrm\{back\}\}\(v\)\+25f\_\{\\mathrm\{aff\}\}\(v\)\+20f\_\{\\mathrm\{fault\}\}\(v\)\+0\.10​fdown​\(v\)\+0\.05​ftype​\(v\),\\displaystyle\+10f\_\{\\mathrm\{down\}\}\(v\)\+05f\_\{\\mathrm\{type\}\}\(v\),wherefback=1−d/dmaxf\_\{\\mathrm\{back\}\}=1\-d/d\_\{\\max\}measures backward relevance,fafff\_\{\\mathrm\{aff\}\}andffaultf\_\{\\mathrm\{fault\}\}indicate affected region membership and direct faulty memory involvement,fdown=min⁡\(1,reached​\_​steps/8\)f\_\{\\mathrm\{down\}\}=\\min\(1,\\mathrm\{reached\\\_steps\}/8\)measures downstream reachability and is raised to at least0\.50\.5if the step reaches the erroneous final answer, andftypef\_\{\\mathrm\{type\}\}is a fixed step type prior set to0\.950\.95,0\.850\.85,0\.750\.75,0\.650\.65, and0\.400\.40for memory read, claim, plan, tool action, and final answer steps, respectively\. The weights and priors are fixed across all domains and fault types\.

If no valid candidate is found, the baseline falls back to the final answer step\. The replay region combines provenance source steps for diagnosed faulty memories with the sequential suffix from the selected root cause step through the erroneous final answer\. Persistent memory cleanup is disabled, so existing faulty memories remain active unless overwritten as a side effect of replay\.

## Appendix IAdapted LongMemEval\-V2 Subset Details

##### Subset Construction\.

We adapt a subset of LongMemEval\-V2\(Wuet al\.[2026](https://arxiv.org/html/2608.10502#bib.bib35)\), which consists of long\-context user\-assistant interaction trajectories\. We do not use the full benchmark distribution\. Instead, we retain only cases for which our pipeline can construct reliable clean data for memory repair evaluation\. Specifically, the trajectory must contain sufficient textual evidence to recover the clean answer, the evidence must be convertible into our controlled benchmark schema, and the clean answer must be verifiable without relying on hidden gold labels\.

This selection criterion favors procedural and navigation\-style tasks, whose answers are grounded in explicit trajectory evidence such as page titles, menu labels, form fields, or ordered action sequences\. We exclude cases that cannot be reliably converted into clean instances, including ambiguous boolean questions, generic short\-answer questions, visual\-only UI states, and aggregation questions whose required evidence is not explicit in the retained trajectory text\. The resulting adapted set contains 50 cases\.

##### Schema Adaptation\.

Each selected trajectory is converted into the same schema used by our controlled benchmark\. Trajectory observations, action traces, page states, and supporting evidence snippets are mapped into memory records with provenance metadata\. Fault construction follows the poisoned memory procedure described in Appendix[D](https://arxiv.org/html/2608.10502#A4)\.

Final answer evaluation follows the same clean answer validation protocol as our main benchmark\. When an answer contains multiple fields or ordered steps, we evaluate it using claim\-level checks derived from the clean trajectory evidence\.

##### Fault Distribution\.

Table[21](https://arxiv.org/html/2608.10502#A9.T21)shows the number of faulty memories per adapted case\. Most cases in this subset contain multiple faulty memories, complementing our controlled benchmark, where most cases contain a single faulty memory\.

Table 21:Fault multiplicity in the adapted LongMemEval\-V2 subset\.
##### Limitations\.

The adapted LongMemEval\-V2 subset is intended as a transfer stress test rather than a replacement for the full LongMemEval\-V2 benchmark\. Because we retain only cases from which reliable clean data can be constructed under our repair schema, the subset does not represent the full LongMemEval\-V2 distribution and should not be used to claim general LongMemEval\-V2 performance\. Its purpose is to test whether repair behavior transfers from our controlled benchmark to trajectory\-derived memory records\.

Similar Articles

Causal Episodic Memory for Feedback-Driven Agent Repair

arXiv cs.CL

This paper introduces MERIT, a training-free agent that uses causal episodic memory of past repair outcomes to improve subsequent Text-to-SQL generations, boosting execution accuracy on Spider and BIRD benchmarks.

Honest Lying: Understanding Memory Confabulation in Reflexive Agents

Hugging Face Daily Papers

This paper identifies memory confabulation in Reflexion-style agents, where agents store incorrect task interpretations and persist in errors across environment resets. The authors introduce the Reflection Repetition Rate (RRR) metric to detect this and propose a mitigation that replaces open-ended self-diagnosis with programmatic failure signal extraction.