Probing Stability-Plasticity Tradeoffs in Agent Memory through Cognitive Experimental Paradigms
Summary
This paper introduces MemProbe, a cognitive-science-inspired framework for evaluating stability-plasticity tradeoffs in agent memory systems through experimental paradigms, providing interpretable profiles of memory maintenance over time.
View Cached Full Text
Cached at: 09/28/26, 09:39 AM
# Probing Stability–Plasticity Tradeoffs in Agent Memory through Cognitive Experimental Paradigms
Source: [https://arxiv.org/html/2609.30558](https://arxiv.org/html/2609.30558)
Jiaqi DingGuorong WuAffiliation:Department of Computer ScienceAffiliation:Department of PsychiatryAffiliation:UNC\-Chapel HillEmail:[guorong\_wu@med\.unc\.edu](mailto:)
###### Abstract
Agent memory systems are increasingly used to maintain long\-term user preferences, task states and evolving facts, but current evaluations often collapse memory behavior into final\-answer accuracy\. We introduceMemProbe, a cognitive\-science\-inspired framework for diagnosing stability–plasticity tradeoffs in agent memory\. The framework is motivated by a core insight from cognitive memory research: memory is reconstructive and shaped by interference, source reliability, reinforcement, and reactivation\.MemProbeturns this insight into four reusable experimental paradigms \(interference, misinformation, consolidation strength, and reconsolidation window\) that manipulate when a memory should be updated, preserved, or treated as uncertain\. It further decomposes correctness into behavioral profiles that reveal how systems update, preserve, attribute, and temporally organize information\. We instantiate these paradigms in a 56\-episode diagnostic suite and evaluate six incremental memory systems under a unified protocol\. Results show that systems with similar aggregate scores exhibit distinct behavioral profiles\.MemProbeprovides such a diagnostic lens, turning aggregate performance into interpretable profiles of memory maintenance over time\. Code is available at[https://github\.com/jq\-ding/MemProbe](https://github.com/jq-ding/MemProbe)\.
## 1Introduction
Two agent memory systems can achieve identical accuracy but behave very differently when updating memory: one may aggressively overwrite prior memories upon seeing new evidence; another may retain outdated information unless explicitly corrected\. Identical scores can thus mask very different risks for long\-term memory \(Figure[1](https://arxiv.org/html/2609.30558#S1.F1)\)\.
Figure 1:Illustrative depiction of how the same accuracy can mask distinct memory\-update behaviors, motivating behavioral evaluation beyond final\-answer correctness\.This distinction has become increasingly important because modern agent memory systems differ substantially in how they store, retrieve, compress, invalidate, and update information\. Retrieval\-only and Full\-context systems often defer conflict resolution to query time, where a reader model interprets the available evidence\. Other systems maintain memory during ingestion by extracting structured facts, compressing interaction histories into evolving summaries, or applying explicit operations such asAdd,Update,Delete, andNoop\. Therefore, these mechanisms can produce different update tendencies even when their final answers are similar\.
Existing benchmarks have substantially expanded the scope of long\-term memory evaluation\. Long\-context and conversational memory benchmarks such as LongMemEval[Wu et al\. \(2024\)](https://arxiv.org/html/2609.30558#bib.bib1), LoCoMo[Maharana et al\. \(2024\)](https://arxiv.org/html/2609.30558#bib.bib2)and MemBench[Tan et al\. \(2025a\)](https://arxiv.org/html/2609.30558#bib.bib23)evaluate temporal reasoning, knowledge updates, and multi\-session understanding\. More recent memory\-agent benchmarks such as MemoryAgentBench[Hu et al\. \(2026b\)](https://arxiv.org/html/2609.30558#bib.bib3), MemoryStress[Sosa \(2026\)](https://arxiv.org/html/2609.30558#bib.bib24), AMA\-Bench[Zhao et al\. \(2026\)](https://arxiv.org/html/2609.30558#bib.bib4), MemoryArena[He et al\. \(2026\)](https://arxiv.org/html/2609.30558#bib.bib5)and EverMemBench[Hu et al\. \(2026a\)](https://arxiv.org/html/2609.30558#bib.bib33)further move toward incremental memory maintenance, contradiction handling, agentic trajectories, and temporally evolving decisions\. These benchmarks are valuable and have pushed the field beyond simple recall\.
However, their primary outcomes remain task\-level correctness measures such as accuracy, F1, recall or task success rate\. Such measures tell us whether a system produced the correct final answer, but not what memory\-update behavior produced that answer\. A system may answer a contradiction case correctly because it retrieves the right evidence, because it has learned a reliable update policy, or simply because the test distribution favors update\-needed examples\. Conversely, a system with high aggregate accuracy may still be over\-plastic or over\-stable\. As agent memory systems become more diverse, evaluating only final\-answer correctness risks collapsing distinct update profiles into a single score\.
We therefore introduceMemProbe, a cognitive\-science\-inspired framework for behaviorally evaluating stability–plasticity tradeoffs in agent memory systems\. Rather than assuming access to an agent’s internal memory,MemProbecontrols the input sequence and observes the resulting behavior\. This follows the experimental logic of cognitive memory research: when internal memory states cannot be directly observed, controlled behavioral paradigms can reveal systematic memory tendencies\.
MemProbe’s primary contribution is behavioral measurements and a set of reusable paradigm specifications for diagnosing memory\-update dynamics\. We make three contributions:
First, we introduce a paradigm\-specification framework for evaluating agent memory update behavior\.MemProbedefines four diagnostic paradigms: Interference\([Underwood, 1957](https://arxiv.org/html/2609.30558#bib.bib11);[Barnes and Underwood, 1959](https://arxiv.org/html/2609.30558#bib.bib8)\), Misinformation\([Loftus, 1975](https://arxiv.org/html/2609.30558#bib.bib21)\), Consolidation Strength\([Ebbinghaus, 1913](https://arxiv.org/html/2609.30558#bib.bib10);[McClelland et al\., 1995](https://arxiv.org/html/2609.30558#bib.bib6)\), and Reconsolidation Window\([Nader et al\., 2000](https://arxiv.org/html/2609.30558#bib.bib7)\)\. Each paradigm specifies what information is encoded, what perturbation is introduced, what variables are controlled, what probes are asked, and what behavioral signatures are expected\.
Second, we define behavioral metrics that decompose aggregate correctness into multi\-dimensional stability–plasticity profiles\. Based on these metrics,MemProbefurther generates a systematic behavioral profile for each evaluated memory system, summarizing its strengths, weaknesses, dominant failure modes and stability–plasticity tendencies in a structured report\.
Third, we instantiate the paradigms in a compact diagnostic suite and apply them to existing memory systems\. Empirically,MemProbeshows that high\-scoring systems can still differ markedly in how they maintain evolving information, motivating behavioral profiles rather than single\-score evaluation\. At the same time,MemProbeis not limited to this fixed suite: we provide the trial structure, latent specification, controlled variables, probe types, gold\-label requirements, and scoring protocol needed for researchers to construct customized suites in their own domains\.
The rest of the paper is organized as follows: Sec\.[3](https://arxiv.org/html/2609.30558#S3)introduces theMemProbeparadigm specification\. Sec\.[4](https://arxiv.org/html/2609.30558#S4)describes the evaluated systems, input protocols and behavioral metrics\. Sec\.[5](https://arxiv.org/html/2609.30558#S5)reports results and analyzes stability–plasticity profiles, paradigm\-wise behavior and failure modes\.
## 2Related Work
Agent Memory Systems\.Modern agent memory systems\([Park et al\., 2023](https://arxiv.org/html/2609.30558#bib.bib38);[Liu et al\., 2026](https://arxiv.org/html/2609.30558#bib.bib39);[Zhong et al\., 2024](https://arxiv.org/html/2609.30558#bib.bib37);[Packer et al\., 2024](https://arxiv.org/html/2609.30558#bib.bib17);[Fang et al\., 2026](https://arxiv.org/html/2609.30558#bib.bib36);[Tan et al\., 2025b](https://arxiv.org/html/2609.30558#bib.bib35)\)differ substantially in how they store, retrieve, compress, invalidate, and update information\. Extractive systems such as Mem0\([Chhikara et al\., 2025](https://arxiv.org/html/2609.30558#bib.bib12)\)and Zep\([Rasmussen et al\., 2025](https://arxiv.org/html/2609.30558#bib.bib13)\)instead maintain structured memories by writing, merging or invalidating facts during ingestion\. Compression\-based systems like Observational Memory\([Barnes, 2026](https://arxiv.org/html/2609.30558#bib.bib27)\)update an implicit summary or observation state through repeated summarization\. Policy\-based systems such as Memory\-R1[Yan et al\. \(2026\)](https://arxiv.org/html/2609.30558#bib.bib18)and AgeMem\([Yu et al\., 2026](https://arxiv.org/html/2609.30558#bib.bib14)\)learn explicit memory operations\. These mechanisms can produce different update tendencies even when their final answers are similar, motivating evaluation that looks beyond task accuracy to the underlying memory behavior\.
Agent Memory Benchmarks\.Recent benchmarks\([Hu et al\., 2026a](https://arxiv.org/html/2609.30558#bib.bib33);[Wang et al\., 2026](https://arxiv.org/html/2609.30558#bib.bib26);[Bian et al\., 2026](https://arxiv.org/html/2609.30558#bib.bib25)\)evaluate how LLMs and agents retain and use information across long interaction histories\. LoCoMo\([Maharana et al\., 2024](https://arxiv.org/html/2609.30558#bib.bib2)\)and LongMemEval\([Wu et al\., 2024](https://arxiv.org/html/2609.30558#bib.bib1)\)probe long\-horizon conversational recall through multi\-session dialogue tasks, while MemoryAgentBench\([Hu et al\., 2026b](https://arxiv.org/html/2609.30558#bib.bib3)\)and MemBench\([Tan et al\., 2025a](https://arxiv.org/html/2609.30558#bib.bib23)\)move toward incremental memory maintenance, treating memory as something accumulated through interaction rather than read from a static long context\.
Another group of benchmarks stresses memory systems with contradiction, noise, and evolving facts\. MemoryStress simulates 1,000 sessions with contradictions, fading memories, and accumulated noise\([Sosa, 2026](https://arxiv.org/html/2609.30558#bib.bib24)\), and MemoryArena evaluates interdependent multi\-session tasks where later success depends on distilling earlier actions and feedback\([He et al\., 2026](https://arxiv.org/html/2609.30558#bib.bib5)\)\. Concurrent with our work, MemConflict\([Tao et al\., 2026](https://arxiv.org/html/2609.30558#bib.bib28)\)examines whether long\-term memory systems retrieve and apply memories that are temporally valid, factually correct, and contextually applicable under conflicting evidence\.MemProbeis complementary, rather than scoring task accuracy under long horizons or conflict, it treats conflict as one manifestation of a broader stability–plasticity trade\-off and uses cognitive\-science\-inspired paradigms to diagnose memory update behaviors\.
Figure 2:Overview of theMemProbeevaluation framework\. Cognitive\-science\-inspired paradigms generate controlled trial episodes that are processed by a memory system and queried through multi\-dimensional probes\. Observable responses are converted into behavioral stability–plasticity profiles\.
## 3From Cognitive Memory Paradigms to Agent Memory Diagnosis
MemProbeformulates evaluation as a behavioral diagnosis problem: controlled contexts manipulate memory conditions, and targeted probes reveal patterns of update behaviors\.
### 3\.1Diagnostic Paradigms from Cognitive Memory Research
Many agent memory systems differ in how they store, retrieve, compress or update information\. Therefore, following the logic of cognitive memory experiments,MemProbediagnoses memory behavior through observable responses rather than implementation\-specific memory states \(Figure[2](https://arxiv.org/html/2609.30558#S2.F2)\)\.
#### Paradigm I: Interference\.
This paradigm is motivated by classicAB−ACAB\-ACpaired\-associate learning studies[Underwood \(1957\)](https://arxiv.org/html/2609.30558#bib.bib11);[Barnes and Underwood \(1959\)](https://arxiv.org/html/2609.30558#bib.bib8)\. Participants first learn an associationAA–BBand later learn a competing associationAA–CC\. When probed withAA, producing the old valueBBafter learningCCreflects proactive interference, whereas producingCCwhen asked for the earlier association reflects retroactive interference\. The paradigm therefore separates successful updating from preservation of historical memory\.
In agent memory, the same logic captures competition among an established memory, a legitimate update and a later distractor\. This paradigm probes excessive stability, where valid updates fail, and excessive plasticity, where distractors intrude as new memories\.
#### Paradigm II: Misinformation\.
This paradigm studies how post\-event misleading information can distort memory for an original event[Loftus and Palmer \(1974\)](https://arxiv.org/html/2609.30558#bib.bib9);[Loftus \(1975\)](https://arxiv.org/html/2609.30558#bib.bib21)\. Its key insight is that later information affects memory not only through its content, but also through its source and credibility: misleading details from trusted sources are more likely to be incorporated, whereas unreliable sources are more likely to be discounted\.
In agent memory, conflicting values may come from users, assistants, tools, retrieved documents, stale trackers or third\-party statements\. The central question is whether the system distinguishes evidence from mere mention\. A reliable user correction or trusted tool update may warrant revision, while a hallucinated assistant statement, rumor or non\-target value should not overwrite memory\. This paradigm tests source\-sensitive updating and misinformation\-driven over\-updating\.
#### Paradigm III: Consolidation Strength\.
This paradigm reflects that strongly reinforced memories are more resistant to disruption\. Overlearning improves retention[Ebbinghaus \(1913\)](https://arxiv.org/html/2609.30558#bib.bib10), and interference studies suggest that repeatedly learned associations are harder to overwrite than weakly learned ones[Underwood \(1957\)](https://arxiv.org/html/2609.30558#bib.bib11)\. Related reconsolidation work further shows that more consolidated memories can be less susceptible to modification even after reactivation[Milekic and Alberini \(2002\)](https://arxiv.org/html/2609.30558#bib.bib22)\.
For agent memory, this means update behavior should depend on prior support\. A fact repeatedly confirmed across interactions should be harder to overwrite than a one\-off statement, especially under weak or unreliable challenges\. However, strong reliable evidence should still allow revision\. This paradigm tests whether a system treats all memories as equally plastic or adjusts stability and plasticity according to memory strength\.
#### Paradigm IV: Reconsolidation Window\.
This paradigm shows that recalling a consolidated memory can temporarily make it modifiable[Nader et al\. \(2000\)](https://arxiv.org/html/2609.30558#bib.bib7);[Schiller et al\. \(2010\)](https://arxiv.org/html/2609.30558#bib.bib19)\. After reactivation, new information may update or distort the original memory before it stabilizes again\. This effect can also be selective: directly reactivated memories are more susceptible to modification than indirectly activated ones[Deębiec et al\. \(2006\)](https://arxiv.org/html/2609.30558#bib.bib20)\.
In agent memory, an old memory may be explicitly recalled before new evidence appears\. Reactivation can help compare old and new information, but it should not make the system indiscriminately accept weak speculation or misinformation\. This paradigm tests whether reactivation changes update sensitivity, and whether that sensitivity is selective to reliable evidence rather than causing reactivation\-induced over\-updating\.
### 3\.2Controlled Trial Design
EachMemProbetrial is an ordered episode in theencoding–perturbation–probingstructure designed to expose a specific memory\-update behavior\. In anencodingstage, an initial target memory is established\. It may then include filler interactions, near\-miss facts or distractors to create a realistic memory context and prevent the task from reducing to simple recency matching\.
The central manipulation occurs in theperturbationstage\. Depending on the paradigm, the perturbation may introduce a valid update, a conflicting but unreliable value, a weak challenge to an established memory or an explicit reactivation of a prior memory\. Some trials also include post\-perturbation context, which tests whether the system is overly sensitive to surface similarity, later mentions or irrelevant recency cues\.
Finally, the episode ends with targetedprobes\. These probes query the system about the current state, previous state, source of information, conflict status or temporal update sequence\. What must be preserved is the controlled relation among the initial memory, the perturbation, the expected behavior and the probe\. App\.[B](https://arxiv.org/html/2609.30558#A2)provides the construction details of our demonstration suite\.
## 4MemProbeEvaluation Protocol
We now specify howMemProbetrials are evaluated in practice\. The goal is not to collapse systems into a single ranking, but to make their responses comparable through a shared input protocol and a common set of behavioral measurements\.
### 4\.1Systems and Input Protocol
MemProbesupports several input protocols \(Figure[3](https://arxiv.org/html/2609.30558#S4.F3)\)\. The first set of input protocols are used as references/baselines \(yellow boxes in Figure[3](https://arxiv.org/html/2609.30558#S4.F3)\)\. AFull\-context readerreceives the complete episode at probe time and serves as an upper\-reference for whether the trial is answerable under complete evidence\. Retrieval\-based baselines[Lewis et al\. \(2020\)](https://arxiv.org/html/2609.30558#bib.bib31), includingNaive RAG\(retrieves the top\-kksessions independently for each probe\),Time\-aware RAG\(additionally incorporates session order and temporal metadata\) andOracle RAG\(provides gold\-relevant evidence sessions to distinguish retrieval failure from reader reasoning failure\), index sessions as retrievable units and answer each probe from retrieved evidence\. We also includeheuristic baselinesthat use simple rule\-based candidate selection, such as preferring the most recent value or selecting values from specific source types\. These heuristics are not intended as memory systems, but serve as sanity checks for shortcut vulnerability\. App\.[E](https://arxiv.org/html/2609.30558#A5)reports schema checks, probe\-level validation, heuristic baselines, and cross\-model validation results\.
Figure 3:Input protocols used inMemProbe\. Full\-context reading, retrieval\-based reading, heuristic baselines, batch memory, and incremental memory correspond to different computational interpretations of memory and are therefore reported separately\. Incremental memory is the primary setting for evaluating stability–plasticity behavior\.Our primary evaluation setting isincremental memory maintenance\(purple boxes in Figure[3](https://arxiv.org/html/2609.30558#S4.F3)\) because it requires the system to decide during ingestion whether new evidence should update, preserve or coexist with prior memory\. For each episode, the memory store is reset and sessions are provided in chronological order\. At probe time, the system must answer using its maintained memory state, retrieval interface or internal memory mechanism\. Formally, for an episodeeewith sessions\{st\}t=1T\\\{s\_\{t\}\\\}\_\{t=1\}^\{T\}and probes\{qj\}j=1m\\\{q\_\{j\}\\\}\_\{j=1\}^\{m\}, the protocol is:
M0\\displaystyle M\_\{0\}←Reset\(\),\\displaystyle\\leftarrow\\mathrm\{Reset\}\(\),\(1\)Mt\\displaystyle M\_\{t\}←Ingest\(Mt−1,st\),\\displaystyle\\leftarrow\\mathrm\{Ingest\}\(M\_\{t\-1\},s\_\{t\}\),t=1,…,T,\\displaystyle t=1,\\ldots,T,a^j\\displaystyle\\hat\{a\}\_\{j\}←Query\(MT,qj\),\\displaystyle\\leftarrow\\mathrm\{Query\}\(M\_\{T\},q\_\{j\}\),j=1,…,m\.\\displaystyle j=1,\\ldots,m\.We evaluate six incremental memory systems under the same sequential\-ingestion protocol:Mem0[Chhikara et al\. \(2025\)](https://arxiv.org/html/2609.30558#bib.bib12),Zep/Graphiti[Rasmussen et al\. \(2025\)](https://arxiv.org/html/2609.30558#bib.bib13),LangMem[The LangChain Team \(2025\)](https://arxiv.org/html/2609.30558#bib.bib15),Cognee[Markovic et al\. \(2025\)](https://arxiv.org/html/2609.30558#bib.bib16),A\-MEM[Xu et al\. \(2025\)](https://arxiv.org/html/2609.30558#bib.bib29), andMemoryOS[Kang et al\. \(2025\)](https://arxiv.org/html/2609.30558#bib.bib30)\. These systems cover extraction\-based memory, temporal knowledge\-graph memory, framework\-level memory primitives, graph\-vector hybrid memory, agentic self\-organizing memory and hierarchical memory\-OS designs\. For consistency, all systems are fed the same ordered sessions and are probed using a unified reader over their retrieved or extracted memory\. When a system requires an internal LLM for memory extraction or consolidation, we use the same internal model whenever the interface permits\.
### 4\.2Behavioral Metrics
All metrics are computed from structured probe scores\. A probe score records whether the system’s answer matches the structured gold fields required by that probe, such as the current/previous values\. We define the full probe taxonomy and scoring procedure in App\.[C](https://arxiv.org/html/2609.30558#A3)\. Let𝒬\\mathcal\{Q\}denote the set of all probes and letc\(q,S\)c\(q,S\)be an indicator of whether systemSSanswers probeqqcorrectly under this scoring procedure\. We define accuracy over any probe subset𝒟\\mathcal\{D\}as:
A\(S,𝒟\)=1\|𝒟\|∑q∈𝒟c\(q,S\)\.A\(S;\\mathcal\{D\}\)=\\frac\{1\}\{\|\\mathcal\{D\}\|\}\\sum\_\{q\\in\\mathcal\{D\}\}c\(q,S\)\.\(2\)Overall accuracy is computed by applying Eq\. \([2](https://arxiv.org/html/2609.30558#S4.E2)\) to all probes:Acc\(S\)=A\(S,𝒬\)\.\\mathrm\{Acc\}\(S\)=A\(S;\\mathcal\{Q\}\)\.
Unless noted otherwise, confidence intervals are episode\-level bootstrap intervals \(2,000 resamples, fixed seed\), and system comparisons use paired bootstrap on the same resampled episodes\.
#### Stability–plasticity metrics\.
Let𝒰\\mathcal\{U\}denote probes from trials where the expected behavior is to accept a valid update, and let𝒫\\mathcal\{P\}denote probes from trials where the expected behavior is to preserve the existing memory\. Using Eq\. \([2](https://arxiv.org/html/2609.30558#S4.E2)\), we define plasticity and stability as:
Plasticity\(S\)\\displaystyle\\mathrm\{Plasticity\}\(S\)=A\(S,𝒰\),\\displaystyle=A\(S;\\mathcal\{U\}\),\(3\)Stability\(S\)\\displaystyle\\mathrm\{Stability\}\(S\)=A\(S,𝒫\)\.\\displaystyle=A\(S;\\mathcal\{P\}\)\.The complementary condition\-level error rates are:
UpdateError\(S\)\\displaystyle\\mathrm\{UpdateError\}\(S\)=1−Plasticity\(S\),\\displaystyle=1\-\\mathrm\{Plasticity\}\(S\),\(4\)PreserveError\(S\)\\displaystyle\\mathrm\{PreserveError\}\(S\)=1−Stability\(S\)\.\\displaystyle=1\-\\mathrm\{Stability\}\(S\)\.To provide a compact summary across the two conditions, we report their harmonic mean:
SP\-Balance\(S\)=2Plasticity\(S\)Stability\(S\)Plasticity\(S\)\+Stability\(S\)\.\\mathrm\{SP\\text\{\-\}Balance\}\(S\)=\\frac\{2\\,\\mathrm\{Plasticity\}\(S\)\\,\\mathrm\{Stability\}\(S\)\}\{\\mathrm\{Plasticity\}\(S\)\+\\mathrm\{Stability\}\(S\)\}\.\(5\)Eq\. \([5](https://arxiv.org/html/2609.30558#S4.E5)\) penalizes systems that perform well on only one side of the stability–plasticity tradeoff\.
#### Historical, source and temporal fidelity\.
We derive fidelity metrics from the corresponding probe types defined in App\.[C\.1](https://arxiv.org/html/2609.30558#A3.SS1)\. Historical fidelity is measured by previous\-value accuracy, source fidelity by source and conflict\-status accuracy, and temporal fidelity by temporal\-probe accuracy\. These metrics measure whether a system preserves not only the current answer, but also the prior state, provenance and update sequence behind that answer\.
MetricOracle RAG\(diag\.\)Full\-contextMem0GraphitiLangMemCogneeA\-MEMMemoryOSNaiveRAGTime\-awareRAGOverall\[95% CI\]93\.3\[89\.7, 96\.4\]88\.8\[84\.4, 92\.9\]41\.5\[35\.7, 46\.9\]62\.5\[55\.8, 68\.8\]65\.2\[57\.6, 72\.3\]81\.7\[76\.3, 86\.6\]88\.4\[83\.9, 92\.4\]60\.3\[53\.6, 66\.5\]68\.3\[60\.3, 76\.3\]68\.8\[61\.2, 75\.9\]
Table 1:Overall behavioral performance across references, incremental memory systems and retrieval baselines\.
#### Retrieval and evidence\-integration metrics\.
For retrieval\-based diagnostic baselines, we compare retrieved\-context performance against full\-context reading\. LetAccFC\\mathrm\{Acc\}\_\{\\mathrm\{FC\}\}denote full\-context accuracy and letAccr\\mathrm\{Acc\}\_\{r\}denote the accuracy of a retrieval\-based conditionrr\. We defineΔFC\(r\)=AccFC−Accr,\\Delta\_\{\\mathrm\{FC\}\}\(r\)=\\mathrm\{Acc\}\_\{\\mathrm\{FC\}\}\-\\mathrm\{Acc\}\_\{r\},andRAGnorm\(r\)=AccrAccFC\.\\mathrm\{RAG\}\_\{\\mathrm\{norm\}\}\(r\)=\\frac\{\\mathrm\{Acc\}\_\{r\}\}\{\\mathrm\{Acc\}\_\{\\mathrm\{FC\}\}\}\.Here,ΔFC\\Delta\_\{\\mathrm\{FC\}\}measures the performance loss relative to full evidence, whileRAGnorm\\mathrm\{RAG\}\_\{\\mathrm\{norm\}\}measures the fraction of full\-context performance recovered by retrieved evidence\. Values above 1 are possible for diagnostic evidence\-selection baselines such as Oracle RAG, because they remove filler and near\-miss sessions and provide only clean evidence\.
For split\-perturbation episodes, where relevant evidence is distributed across sessions, we defineΔsplit=Accsplit−Accnon\-split\.\\Delta\_\{\\mathrm\{split\}\}=\\mathrm\{Acc\}\_\{\\mathrm\{split\}\}\-\\mathrm\{Acc\}\_\{\\mathrm\{non\\text\{\-\}split\}\}\.Negative values therefore indicate a split penalty, while positive values indicate better performance on split episodes\.
#### Failure attribution\.
Incorrect predictions are assigned to diagnostic failure categories based on the structured scores described in App\.[C\.2](https://arxiv.org/html/2609.30558#A3.SS2)\. Retrieval\-based systems may fail because relevant evidence is missing, only part of a split perturbation is retrieved, or the reader misinterprets retrieved evidence\. Incremental systems may fail by missing valid updates, over\-updating on weak evidence, overwriting historical values, losing source information, or retrieving the wrong memory at query time\. These categories are reported separately from aggregate accuracy\.
## 5Results and Analysis
All incremental memory systems are evaluated with a shared protocol: each episode uses a fresh memory store, sessions are ingested sequentially, retrieved/extracted memory is passed to a unified Gemini\-3\-Flash reader, and answers are scored using the same structured JSON scorer\. §[5\.4](https://arxiv.org/html/2609.30558#S5.SS4)verifies that the resulting profiles do not depend on which reader is used\. For systems requiring an internal LLM for extraction, we use Gemini\-2\.5\-Flash when configurable\. The LLM\-selection rationale and full implementation details are provided in App\.[A](https://arxiv.org/html/2609.30558#A1)\.
### 5\.1Overall Memory\-System Performance
Table[1](https://arxiv.org/html/2609.30558#S4.T1)provides a coarse entry point into system behavior\. We report overall accuracy to summarize performance, but do not treat it as the final evaluation criterion\. We separate diagnostic references from real memory systems\. Among incremental memory systems, A\-MEM performs best, reaching 88\.4% overall accuracy and nearly matching the Full\-context upper reference at 88\.8%\. Cognee ranks second with 81\.7% overall accuracy, followed by LangMem, Graphiti, and MemoryOS in the middle range\.
Figure 4:Probe\-level performance across systems\. Many systems perform well on current\-value, previous\-value and change\-detection probes, but degrade on source, conflict\-value and temporal probes\.Figure[4](https://arxiv.org/html/2609.30558#S5.F4)further decomposes performance by probe type \(Probe definitions are provided in App\.[C\.1](https://arxiv.org/html/2609.30558#A3.SS1)\)\. Most systems perform better on current\-value, previous\-value, and change\-detection probes than on source, conflict\-value, and temporal probes\. This suggests that many memory systems can retain or retrieve factual values, but often fail to preserve provenance and update\-sequence information\.
These aggregate scores provide only a coarse summary\. As the following sections show, systems with similar overall accuracy can differ sharply in their update tendencies, historical retention, provenance preservation, temporal reasoning, and robustness to split evidence\. We therefore decompose the results into stability–plasticity profiles, paradigm\-wise behavior, and failure modes below\.
### 5\.2Stability–Plasticity Profiles
Aggregate accuracy alone cannot distinguish systems that update too aggressively from systems that preserve outdated memories too strongly\. We therefore compare systems by plasticity, stability and the corresponding condition\-level error rates\. Plasticity measures whether a system accepts valid updates, while stability measures whether it resists weak or unreliable evidence\.
SystemPlast\.Stab\.SP\-BalPreserveErr\.UpdateErr\.Mem028\.654\.537\.545\.571\.4Graphiti58\.966\.162\.333\.941\.1LangMem63\.467\.065\.133\.036\.6Cognee75\.088\.481\.111\.625\.0A\-MEM91\.185\.788\.314\.38\.9MemoryOS54\.566\.159\.733\.945\.5
Table 2:Stability–plasticity metrics for incremental memory systems\.Figure 5:Stability–plasticity plane\. Systems closer to the lower\-left corner have better SP\-balance\.SystemInterferenceMisinformationConsolidationReconsolidationOverallOracle RAG \(diag\.\)89\.392\.998\.292\.993\.3Full\-context82\.191\.191\.191\.188\.8Mem041\.146\.433\.944\.641\.5Graphiti53\.660\.775\.060\.762\.5LangMem60\.766\.169\.664\.365\.2Cognee80\.480\.485\.780\.481\.7A\-MEM82\.185\.796\.489\.388\.4MemoryOS66\.155\.457\.162\.560\.3Naive RAG67\.975\.060\.769\.668\.3Time\-aware RAG67\.973\.266\.167\.968\.8Table 3:Paradigm\-wise accuracy across the four paradigms\. Interference is difficult even for strong references, while consolidation sharply separates structured memory systems from flatter extraction or retrieval systems\.Table[2](https://arxiv.org/html/2609.30558#S5.T2)reports the stability–plasticity metrics\. A\-MEM has the best balance among memory systems, with 91\.1% plasticity, 85\.7% stability, and an SP\-Balance of 88\.3%\. Cognee is more stability\-biased: it achieves the highest stability at 88\.4%, but lower plasticity at 75\.0%\. LangMem and Graphiti occupy a middle region with different behaviors: LangMem consolidates aggressively and loses history, while Graphiti retains historical values but struggles to update and preserve source information\.
Figure[5](https://arxiv.org/html/2609.30558#S5.F5)plots systems in the stability–plasticity plane\. The ideal region is the lower\-left corner, where errors are low under both conditions\. A\-MEM is closest to this region\. Cognee is also strong but more conservative, with fewer errors when preservation is required and more when an update is required\. Mem0 lies in the least desirable region\.
The three mid\-range systems \(Graphiti, LangMem, MemoryOS\) are statistically indistinguishable in overall accuracy \(all pairwisep\>0\.37p\>0\.37\), and thus directly fall into the “same accuracy” condition\. But their behavior differs once correctness is decomposed by probe type\. Graphiti preserves historical values far better than LangMem \(previous\-value96\.496\.4vs\.67\.967\.9, 95% CI\[\+9\.1,\+48\.3\]\[\+9\.1,\+48\.3\],p=0\.002p=0\.002\), a gap that survives Bonferroni correction across all 18 system–probe comparisons, while the current\-value gap runs in the opposite direction \(66\.166\.1vs\.80\.480\.4,Δ=−14\.3\\Delta=\-14\.3pp\)\. Two systems with the same aggregate score therefore realize opposite maintenance strategies: one biased toward retaining the past, the other toward tracking the present\.
These results illustrate the central motivation ofMemProbe: final accuracy alone cannot reveal whether a memory system is over\-plastic, over\-stable or balanced\.
### 5\.3Paradigm\-wise Analysis
We next analyze system behavior across the fourMemProbeparadigms\. Each paradigm contains 14 episodes and 56 scored probes\. Table[3](https://arxiv.org/html/2609.30558#S5.T3)reports paradigm\-wise accuracy\.
Several patterns emerge\.First, interference is the hardest paradigm for the strongest references: Full\-context reaches 82\.1%, A\-MEM 82\.1%, and Oracle RAG 89\.3%\. This suggests that competing similar facts stress memory behavior even when clean evidence is available\.Second, consolidation sharply separates systems\. A\-MEM reaches 96\.4%, and Graphiti achieves its best paradigm score at 75\.0%, suggesting that note\-based and graph\-based structures help retain accumulated evidence\. Mem0 performs less well on consolidation in our setting, suggesting that repeated reinforcement is not always preserved by flat extracted\-memory representations\.Third, MemoryOS shows a different pattern from most other systems: it performs best on interference but worst on misinformation and consolidation\. This suggests that its heat\-promotion mechanism may keep recent or competing items accessible while failing to preserve reinforced history\.Finally, misinformation favors systems that preserve raw or well\-attributed evidence\. Retrieval\-style access performs relatively well, A\-MEM and Cognee remain strong\. In contrast, systems that aggressively extract, merge or rewrite memories often blur provenance during memory construction\. App\.[D](https://arxiv.org/html/2609.30558#A4)further reports how much independent signal each decomposition axis carries\.
### 5\.4Reader Ablation
To check whether the profiles reflect the memory representation rather than the reader’s ability to interpret it, we re\-ran all six memory systems on the same stored memories with a second reader, Qwen3\-32B, drawn from a different model family than both the main reader \(Gemini\) and the generator \(GPT\) used for dialogue surface form\. So any difference is attributable to the reader alone\.
SystemGeminiQwen3A\-MEM88\.472\.8Cognee81\.768\.3LangMem65\.260\.3Graphiti62\.561\.6MemoryOS60\.351\.8Mem041\.541\.5Table 4:Reader ablation\. The same stored memories are answered by two readers from different model families\. Absolute scores drop under the weaker reader, but the ordering is largely preserved\.Absolute scores drop under Qwen3\-32B, which is expected given its weaker source and conflict\-status accuracy \(Table[4](https://arxiv.org/html/2609.30558#S5.T4)\)\. However, the ordering is essentially unchanged \(Spearmanρ=0\.94\\rho=0\.94, Kendallτ=0\.87\\tau=0\.87\)\. The single exception is LangMem and Graphiti, which swap positions, whose aggregate scores are statistically indistinguishable \(§[5\.2](https://arxiv.org/html/2609.30558#S5.SS2)\), and whose order is therefore not expected to be stable under any perturbation\.
A reader with weaker provenance reasoning lowers every system’s score, but it does not change which system retains more provenance than another\. The behavioral profiles reported here are therefore properties of the maintained memory rather than artifacts of the selected reader\. Full cross\-reader results are given in App\.[E\.5](https://arxiv.org/html/2609.30558#A5.SS5)\.
### 5\.5Failure Taxonomy
MemProbeis designed to identify not only whether a system fails, but how it fails\. We therefore group errors into several recurring failure modes: provenance loss, temporal reconstruction failure, distributed\-evidence retrieval miss, historical overwrite, under\-update, over\-update, and extraction fragility\.
SystemHist\.SourceTemp\.Oracle RAG \(diag\.\)100\.075\.877\.0Full\-context100\.066\.177\.0Mem071\.017\.736\.0Graphiti \(Zep\)96\.041\.945\.0LangMem68\.046\.841\.0Cognee96\.064\.573\.0A\-MEM100\.064\.568\.0MemoryOS89\.033\.955\.0Naive RAG75\.064\.545\.0Time\-aware RAG86\.059\.745\.0Table 5:Historical, source and temporal fidelity across systemsThe sharpest systematic failure is provenance loss\. This indicates that these memory systems can store values but often fail to retain who introduced them, whether they were accepted and whether they applied to the target fact\. This could in principle reflect a reader that fails to recover provenance rather than a memory that fails to retain it\. Two observations separate the two\. First, as shown in Table[5](https://arxiv.org/html/2609.30558#S5.T5), Full\-context uses the same reader on complete evidence and reaches 66\.1%, so part of the difficulty is inherent to source recovery\. But systems fall far below this same\-reader ceiling \(Graphiti 41\.9%, Mem0 17\.7%\), and a gap of this size under an identical reader cannot be a read\-out artifact\. Second, we manually inspected every source error and labeled it as a storage\-side loss \(provenance absent from the retrieved memory, unrecoverable by any reader\) or a read\-out failure \(provenance present but not recovered\)\. As shown in Table[6](https://arxiv.org/html/2609.30558#S5.T6), read\-out errors are roughly constant across systems \(6–9\), so the reader does not explain the large differences in source fidelity\. What varies is storage\-side loss, which grows as systems retain less of the original context\.
SystemRead\-outStorage\-sideA\-MEM62Cognee91LangMem910Graphiti811MemoryOS913Mem0720Table 6:Manual classification of source\-probe errors\. Systems are ordered by how much of the original context they preserve for the reader\. Read\-out errors are roughly constant, whereas storage\-side losses grow as systems compress memory more aggressively\.Temporal fidelity is also weak: most memory systems fall below the Full\-context temporal score of 77%, suggesting that they do not reliably reconstruct update order, trigger and reason\.
Another clear failure mode appears in the split\-perturbation setting, where the update trigger/source and the accepted current value are placed in separate sessions\. Solving such episodes requires retrieving and integrating both perturbation parts, rather than finding a single decisive update sentence\. In Table[7](https://arxiv.org/html/2609.30558#S5.T7), Time\-aware RAG adds true temporal position labels to retrieved sessions, but improves only marginally over Naive RAG: 68\.8% overall versus 68\.3%, and 33\.3% split accuracy versus 29\.2%\. Split current\-value \(Split CV\) and change\-detection \(Split CD\) accuracy remain stuck at 8\.3% for both Naive and Time\-aware RAG\. This shows that the problem is not missing temporal order information\.
In contrast, Oracle RAG reaches 93\.3% overall and 89\.6% on split episodes\. Its split current\-value accuracy is 91\.7%, and its split change\-detection accuracy is 100%\. Since Oracle RAG uses the same reader but receives clean non\-filler evidence, the gap between Naive RAG and Oracle RAG isolates retrieval failure as the main bottleneck\. In other words, split failures occur because the retriever often fails to recover both perturbation parts, not because the reader cannot reason over the evidence\.
MetricNaiveRAGTime\-awareRAGOracleRAGFull\-contextΔFC\{\\Delta\_\{\\mathrm\{FC\}\}\}20\.520\.0\-4\.5–RAGnorm\{\\mathrm\{RAG\}\_\{\\mathrm\{norm\}\}\}0\.7690\.7751\.051–Non\-split79\.078\.494\.390\.9Split29\.233\.389\.681\.2Split CV8\.38\.391\.766\.7Split CD8\.38\.3100\.091\.7
Table 7:Diagnostic RAG ablation for split\-perturbation episodes\.ΔFC\\Delta\_\{\\mathrm\{FC\}\}measures the loss relative to full\-context reading, andRAGnorm\\mathrm\{RAG\}\_\{\\mathrm\{norm\}\}measures the fraction of full\-context performance recovered by retrieved evidence\. Split CV denotes current\-value accuracy, and Split CD denotes change\-detection accuracy\.Finally, Figure[6](https://arxiv.org/html/2609.30558#S5.F6)summarizes split penalties across systems\. Most systems degrade when evidence is distributed across sessions\. MemoryOS is the only system with a positive split delta, but this should be interpreted cautiously because the split subset contains fewer source and conflict probes, which are its weakest categories\. Overall, the split results show that distributed evidence integration remains a major weakness for retrieval\-based and incremental memory systems\.
Figure 6:Split penalty/gain by systems\. Negative values indicate that performance drops when perturbation evidence is distributed across sessions\.
## 6Discussion and Conclusion
MemProbeprovides a cognitive\-style diagnostic framework for evaluating agent memory systems beyond aggregate final\-answer accuracy\. By organizing memory evaluation around interference, misinformation, consolidation, and reconsolidation paradigms,MemProbeexposes how systems balance stability and plasticity under controlled update and preservation conditions\.
MemProbeis not intended to be an exhaustive leaderboard\. Our current suite is a compact instantiation of the proposed paradigm specifications, and future work can instantiate the same trial structure in other domains, modalities and real user histories\. Nevertheless, the framework demonstrates that agent memory evaluation should move from asking only whether a system answers correctly to asking how it updates, preserves and distorts information over time\.
## Limitations
MemProbeis a diagnostic framework rather than an exhaustive benchmark\. While our new generated suite demonstrates that the paradigm specifications can be systematically scaled beyond the 56\-episode reference suite, the current evaluation still covers only a limited range of domains, languages, user histories and deployment settings\. We use a unified reader and structured scoring protocol to enable controlled comparison across systems, and additional analysis suggests that the main findings are not driven by this choice, although native end\-to\-end deployments may differ\. Finally, because part of the dialogue surface form is generated with LLM, the suite may still contain generator\-specific biases in wording, scenario realization, or interaction style\.
## References
- Barnes and Underwood \(1959\)J\. M\. Barnes and B\. J\. Underwood" Fate" of first\-list associations in transfer theory\.\.Journal of experimental psychology58\(2\),pp\. 97\.Cited by:[§1](https://arxiv.org/html/2609.30558#S1.p7.1),[§3\.1](https://arxiv.org/html/2609.30558#S3.SS1.SSS0.Px1.p1.1)\.
- Barnes \(2026\)T\. BarnesAnnouncing observational memory\.Note:[https://mastra\.ai/blog/observational\-memory](https://mastra.ai/blog/observational-memory)Mastra Blog\. Accessed: 2026\-05\-24Cited by:[§2](https://arxiv.org/html/2609.30558#S2.p1.1)\.
- Bianet al\.\(2026\)H\. Bian, Z\. Yao, S\. Hu, Z\. Xu, S\. Zhang, Y\. Guo, Z\. Yang, X\. Han, H\. Wang, and R\. ChenRealMem: benchmarking llms in real\-world memory\-driven interaction\.External Links:2601\.06966,[Link](https://arxiv.org/abs/2601.06966)Cited by:[§2](https://arxiv.org/html/2609.30558#S2.p2.1)\.
- Chhikaraet al\.\(2025\)P\. Chhikara, D\. Khant, S\. Aryan, T\. Singh, and D\. YadavMem0: building production\-ready ai agents with scalable long\-term memory\.arXiv preprint arXiv:2504\.19413\.Cited by:[§2](https://arxiv.org/html/2609.30558#S2.p1.1),[§4\.1](https://arxiv.org/html/2609.30558#S4.SS1.p3.1)\.
- Deębiecet al\.\(2006\)J\. Deębiec, V\. Doyère, K\. Nader, and J\. E\. LeDouxDirectly reactivated, but not indirectly reactivated, memories undergo reconsolidation in the amygdala\.Proceedings of the National Academy of Sciences103\(9\),pp\. 3428–3433\.Cited by:[§3\.1](https://arxiv.org/html/2609.30558#S3.SS1.SSS0.Px4.p1.1)\.
- Ebbinghaus \(1913\)H\. EbbinghausMemory: a contribution to experimental psychology\.Teachers College Press\.External Links:[Document](https://dx.doi.org/10.1037/10011-000)Cited by:[§1](https://arxiv.org/html/2609.30558#S1.p7.1),[§3\.1](https://arxiv.org/html/2609.30558#S3.SS1.SSS0.Px3.p1.1)\.
- Fanget al\.\(2026\)J\. Fang, X\. Deng, H\. Xu, Z\. Jiang, Y\. Tang, Z\. Xu, S\. Deng, Y\. Yao, M\. Wang, S\. Qiao, H\. Chen, and N\. ZhangLightmem: lightweight and efficient memory\-augmented generation\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 98706–98729\.Cited by:[§2](https://arxiv.org/html/2609.30558#S2.p1.1)\.
- Heet al\.\(2026\)Z\. He, Y\. Wang, C\. Zhi, Y\. Hu, T\. Chen, L\. Yin, Z\. Chen, T\. A\. Wu, S\. Ouyang, Z\. Wang, J\. Pei, J\. McAuley, Y\. Choi, and A\. PentlandMemoryArena: benchmarking agent memory in interdependent multi\-session agentic tasks\.External Links:2602\.16313,[Link](https://arxiv.org/abs/2602.16313)Cited by:[item 4](https://arxiv.org/html/2609.30558#A2.I1.i4.p1.1),[§1](https://arxiv.org/html/2609.30558#S1.p3.1),[§2](https://arxiv.org/html/2609.30558#S2.p3.1)\.
- Huet al\.\(2026a\)C\. Hu, T\. Li, X\. Gao, H\. Chen, Y\. Bai, D\. Xu, T\. Lin, X\. Li, Y\. Han, J\. Pei, and Y\. DengEvaluating long\-horizon memory for multi\-party collaborative dialogues\.External Links:2602\.01313,[Link](https://arxiv.org/abs/2602.01313)Cited by:[§1](https://arxiv.org/html/2609.30558#S1.p3.1),[§2](https://arxiv.org/html/2609.30558#S2.p2.1)\.
- Huet al\.\(2026b\)Y\. Hu, Y\. Wang, and J\. McAuleyEvaluating memory in llm agents via incremental multi\-turn interactions\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 156259–156291\.Cited by:[§1](https://arxiv.org/html/2609.30558#S1.p3.1),[§2](https://arxiv.org/html/2609.30558#S2.p2.1)\.
- Kanget al\.\(2025\)J\. Kang, M\. Ji, Z\. Zhao, and T\. BaiMemory os of ai agent\.External Links:2506\.06326,[Link](https://arxiv.org/abs/2506.06326)Cited by:[§4\.1](https://arxiv.org/html/2609.30558#S4.SS1.p3.1)\.
- Latimeret al\.\(2025\)C\. Latimer, N\. Boschi, A\. Neeser, C\. Bartholomew, G\. Srivastava, X\. Wang, and N\. RamakrishnanHindsight is 20/20: building agent memory that retains, recalls, and reflects\.External Links:2512\.12818,[Link](https://arxiv.org/abs/2512.12818)Cited by:[2nd item](https://arxiv.org/html/2609.30558#A1.I3.i2.p1.1)\.
- Lewiset al\.\(2020\)P\. Lewis, E\. Perez, A\. Piktus, F\. Petroni, V\. Karpukhin, N\. Goyal, H\. Küttler, M\. Lewis, W\. Yih, T\. Rocktäschel, S\. Riedel, and D\. KielaRetrieval\-augmented generation for knowledge\-intensive nlp tasks\.InProceedings of the 34th International Conference on Neural Information Processing Systems,NIPS ’20,Red Hook, NY, USA\.External Links:ISBN 9781713829546Cited by:[§4\.1](https://arxiv.org/html/2609.30558#S4.SS1.p1.1)\.
- Liuet al\.\(2026\)J\. Liu, Y\. Su, P\. Xia, Y\. Zhou, S\. Han, Z\. Zheng, C\. Xie, M\. Ding, and H\. YaoSimpleMem: efficient lifelong memory for llm agents\.arXiv preprint arXiv:2601\.02553\.External Links:[Link](https://arxiv.org/abs/2601.02553)Cited by:[§2](https://arxiv.org/html/2609.30558#S2.p1.1)\.
- Loftus and Palmer \(1974\)E\. F\. Loftus and J\. C\. PalmerReconstruction of automobile destruction: an example of the interaction between language and memory\.Journal of verbal learning and verbal behavior13\(5\),pp\. 585–589\.Cited by:[§3\.1](https://arxiv.org/html/2609.30558#S3.SS1.SSS0.Px2.p1.1)\.
- Loftus \(1975\)E\. F\. LoftusLeading questions and the eyewitness report\.Cognitive psychology7\(4\),pp\. 560–572\.Cited by:[§1](https://arxiv.org/html/2609.30558#S1.p7.1),[§3\.1](https://arxiv.org/html/2609.30558#S3.SS1.SSS0.Px2.p1.1)\.
- Maharanaet al\.\(2024\)A\. Maharana, D\. Lee, S\. Tulyakov, M\. Bansal, F\. Barbieri, and Y\. FangEvaluating very long\-term conversational memory of llm agents\.arXiv preprint arXiv:2402\.17753\.Cited by:[item 4](https://arxiv.org/html/2609.30558#A2.I1.i4.p1.1),[§1](https://arxiv.org/html/2609.30558#S1.p3.1),[§2](https://arxiv.org/html/2609.30558#S2.p2.1)\.
- Markovicet al\.\(2025\)V\. Markovic, L\. Obradovic, L\. Hajdu, and J\. PavlovicOptimizing the interface between knowledge graphs and llms for complex reasoning\.External Links:2505\.24478,[Link](https://arxiv.org/abs/2505.24478)Cited by:[§4\.1](https://arxiv.org/html/2609.30558#S4.SS1.p3.1)\.
- McClellandet al\.\(1995\)J\. L\. McClelland, B\. L\. McNaughton, and R\. C\. O’ReillyWhy there are complementary learning systems in the hippocampus and neocortex: insights from the successes and failures of connectionist models of learning and memory\.\.Psychological review102\(3\),pp\. 419\.Cited by:[§1](https://arxiv.org/html/2609.30558#S1.p7.1)\.
- Milekic and Alberini \(2002\)M\. H\. Milekic and C\. M\. AlberiniTemporally graded requirement for protein synthesis following memory reactivation\.Neuron36\(3\),pp\. 521–525\.Cited by:[§3\.1](https://arxiv.org/html/2609.30558#S3.SS1.SSS0.Px3.p1.1)\.
- Naderet al\.\(2000\)K\. Nader, G\. E\. Schafe, and J\. E\. Le DouxFear memories require protein synthesis in the amygdala for reconsolidation after retrieval\.Nature406\(6797\),pp\. 722–726\.Cited by:[§1](https://arxiv.org/html/2609.30558#S1.p7.1),[§3\.1](https://arxiv.org/html/2609.30558#S3.SS1.SSS0.Px4.p1.1)\.
- Packeret al\.\(2024\)C\. Packer, S\. Wooders, K\. Lin, V\. Fang, S\. G\. Patil, I\. Stoica, and J\. E\. GonzalezMemGPT: towards llms as operating systems\.External Links:2310\.08560,[Link](https://arxiv.org/abs/2310.08560)Cited by:[1st item](https://arxiv.org/html/2609.30558#A1.I3.i1.p1.1),[§2](https://arxiv.org/html/2609.30558#S2.p1.1)\.
- Parket al\.\(2023\)J\. S\. Park, J\. C\. O’Brien, C\. J\. Cai, M\. R\. Morris, P\. Liang, and M\. S\. BernsteinGenerative agents: interactive simulacra of human behavior\.InIn the 36th Annual ACM Symposium on User Interface Software and Technology \(UIST ’23\),UIST ’23,New York, NY, USA\.Cited by:[§2](https://arxiv.org/html/2609.30558#S2.p1.1)\.
- Rasmussenet al\.\(2025\)P\. Rasmussen, P\. Paliychuk, T\. Beauvais, J\. Ryan, and D\. ChalefZep: a temporal knowledge graph architecture for agent memory\.External Links:2501\.13956,[Link](https://arxiv.org/abs/2501.13956)Cited by:[§2](https://arxiv.org/html/2609.30558#S2.p1.1),[§4\.1](https://arxiv.org/html/2609.30558#S4.SS1.p3.1)\.
- Sarinet al\.\(2025\)S\. Sarin, L\. Singh, B\. Sarmah, and D\. MehtaMemoria: a scalable agentic memory framework for personalized conversational ai\.External Links:2512\.12686,[Link](https://arxiv.org/abs/2512.12686)Cited by:[4th item](https://arxiv.org/html/2609.30558#A1.I3.i4.p1.1)\.
- Schilleret al\.\(2010\)D\. Schiller, M\. Monfils, C\. M\. Raio, D\. C\. Johnson, J\. E\. LeDoux, and E\. A\. PhelpsPreventing the return of fear in humans using reconsolidation update mechanisms\.Nature463\(7277\),pp\. 49–53\.Cited by:[§3\.1](https://arxiv.org/html/2609.30558#S3.SS1.SSS0.Px4.p1.1)\.
- Sosa \(2026\)J\. SosaWhy i built memorystress\.Note:[https://omegamax\.co/blog/why\-we\-built\-memorystress](https://omegamax.co/blog/why-we-built-memorystress)Accessed: 2026\-05\-13Cited by:[§1](https://arxiv.org/html/2609.30558#S1.p3.1),[§2](https://arxiv.org/html/2609.30558#S2.p3.1)\.
- Tanet al\.\(2025a\)H\. Tan, Z\. Zhang, C\. Ma, X\. Chen, Q\. Dai, and Z\. DongMemBench: towards more comprehensive evaluation on the memory of llm\-based agents\.External Links:2506\.21605,[Link](https://arxiv.org/abs/2506.21605)Cited by:[§1](https://arxiv.org/html/2609.30558#S1.p3.1),[§2](https://arxiv.org/html/2609.30558#S2.p2.1)\.
- Tanet al\.\(2025b\)Z\. Tan, J\. Yan, I\. Hsu, R\. Han, Z\. Wang, L\. Le, Y\. Song, Y\. Chen, H\. Palangi, G\. Lee, A\. R\. Iyer, T\. Chen, H\. Liu, C\. Lee, and T\. PfisterIn prospect and retrospect: reflective memory management for long\-term personalized dialogue agents\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 8416–8439\.External Links:[Link](https://aclanthology.org/2025.acl-long.413/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.413),ISBN 979\-8\-89176\-251\-0Cited by:[§2](https://arxiv.org/html/2609.30558#S2.p1.1)\.
- Taoet al\.\(2026\)Z\. Tao, J\. Zhao, P\. Liu, D\. Xi, Y\. Chen, W\. Xu, and Z\. LiMemConflict: evaluating long\-term memory systems under memory conflicts\.External Links:2605\.20926,[Link](https://arxiv.org/abs/2605.20926)Cited by:[§2](https://arxiv.org/html/2609.30558#S2.p3.1)\.
- The LangChain Team \(2025\)The LangChain TeamLangMem sdk for agent long\-term memory\.Note:[https://www\.langchain\.com/blog/langmem\-sdk\-launch](https://www.langchain.com/blog/langmem-sdk-launch)Accessed: 2026\-05\-13Cited by:[§4\.1](https://arxiv.org/html/2609.30558#S4.SS1.p3.1)\.
- Underwood \(1957\)B\. J\. UnderwoodInterference and forgetting\.\.Psychological review64\(1\),pp\. 49\.Cited by:[§1](https://arxiv.org/html/2609.30558#S1.p7.1),[§3\.1](https://arxiv.org/html/2609.30558#S3.SS1.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2609.30558#S3.SS1.SSS0.Px3.p1.1)\.
- Wanget al\.\(2026\)Y\. Wang, Z\. Zhang, M\. Chi, K\. Yu, Y\. Li, M\. Peng, B\. Tong, C\. Zhang, Y\. Zhou, and J\. LiEvoMemBench: benchmarking agent memory from a self\-evolving perspective\.External Links:2605\.18421,[Link](https://arxiv.org/abs/2605.18421)Cited by:[§2](https://arxiv.org/html/2609.30558#S2.p2.1)\.
- Wuet al\.\(2024\)D\. Wu, H\. Wang, W\. Yu, Y\. Zhang, K\. Chang, and D\. YuLongMemEval: benchmarking chat assistants on long\-term interactive memory\.External Links:2410\.10813,[Link](https://arxiv.org/abs/2410.10813)Cited by:[item 4](https://arxiv.org/html/2609.30558#A2.I1.i4.p1.1),[§1](https://arxiv.org/html/2609.30558#S1.p3.1),[§2](https://arxiv.org/html/2609.30558#S2.p2.1)\.
- Xuet al\.\(2025\)W\. Xu, Z\. Liang, K\. Mei, H\. Gao, J\. Tan, and Y\. ZhangA\-mem: agentic memory for llm agents\.InAdvances in Neural Information Processing Systems,Cited by:[§4\.1](https://arxiv.org/html/2609.30558#S4.SS1.p3.1)\.
- Yanet al\.\(2026\)S\. Yan, X\. Yang, Z\. Huang, E\. Nie, Z\. Ding, Z\. Li, X\. Ma, J\. Bi, K\. Kersting, J\. Z\. Pan, H\. Schütze, V\. Tresp, and Y\. MaMemory\-r1: enhancing large language model agents to manage and utilize memories via reinforcement learning\.External Links:2508\.19828,[Link](https://arxiv.org/abs/2508.19828)Cited by:[3rd item](https://arxiv.org/html/2609.30558#A1.I3.i3.p1.1),[§2](https://arxiv.org/html/2609.30558#S2.p1.1)\.
- Yuet al\.\(2026\)Y\. Yu, L\. Yao, Y\. Xie, Q\. Tan, J\. Feng, Y\. Li, and L\. WuAgentic memory: learning unified long\-term and short\-term memory management for large language model agents\.External Links:2601\.01885,[Link](https://arxiv.org/abs/2601.01885)Cited by:[§2](https://arxiv.org/html/2609.30558#S2.p1.1)\.
- Zhaoet al\.\(2026\)Y\. Zhao, B\. Yuan, J\. Huang, H\. Yuan, Z\. Yu, H\. Xu, L\. Hu, A\. Shankarampeta, Z\. Huang, W\. Ni, Y\. Tian, and J\. ZhaoAMA\-bench: evaluating long\-horizon memory for agentic applications\.External Links:2602\.22769,[Link](https://arxiv.org/abs/2602.22769)Cited by:[§1](https://arxiv.org/html/2609.30558#S1.p3.1)\.
- Zhonget al\.\(2024\)W\. Zhong, L\. Guo, Q\. Gao, H\. Ye, and Y\. WangMemorybank: enhancing large language models with long\-term memory\.InProceedings of the AAAI conference on artificial intelligence,Vol\.38,pp\. 19724–19731\.Cited by:[§2](https://arxiv.org/html/2609.30558#S2.p1.1)\.
## Appendix AImplementation and Reproducibility Details
### A\.1Code and Data Availability
We provide a complete reproducibility package in our Github \(https://github\.com/jq\-ding/MemProbe\)\. The archive includes all code used to generate and validate the demonstration suite, the full 56\-episode dataset and latent specifications, the runner scripts for all evaluated memory systems, raw model outputs, computed metrics, result tables, and analysis scripts\. It also contains detailed implementation notes documenting environment setup, model configurations, system\-specific adaptations, excluded systems, and troubleshooting details\. In addition, we include instructions for constructing customizedMemProbesuites, including how to define new fact types, create latent specifications, generate dialogues, validate gold labels, and evaluate memory systems under the same protocol\.
### A\.2Environment and Model Configuration
The main environment used a unified Gemini\-based evaluation stack\. We usedgemini\-3\-flash\-previewas the final reader for all systems andgemini\-2\.5\-flashas the internal LLM for memory extraction, consolidation, graph construction, or memory promotion whenever a system required an LLM component and allowed this configuration\. When a system required embeddings, we used the system\-supported embedding backend; for Gemini\-compatible systems, we usedgemini\-embedding\-001\. For systems relying on local sentence\-transformer embeddings, we used the default or recommended local embedding model\.
#### Model\-selection rationale\.
We use Gemini\-3\-Flash as the primary reader rather than a GPT\-family model because a substantial portion of the demonstration\-suite dialogue text was generated with GPT\-5\.4 assistance, although the latent specifications, gold labels, and validation checks were manually reviewed\. Using a non\-GPT reader reduces generator and reader coupling in the main evaluation\. We also ran cross\-model validation with GPT\-5\.4, Gemini\-3\-Flash, and Qwen3\-32B \(App\.[E\.5](https://arxiv.org/html/2609.30558#A5.SS5)\)\. The qualitative diagnostic patterns were consistent across models, but Qwen3\-32B was weaker on source attribution and conflict\-status fields, making it less appropriate as the default reader\. We therefore select Gemini\-3\-Flash as a strong independent reader\. For internal memory operations, we use Gemini\-2\.5\-Flash when configurable because it is supported by most evaluated systems, provides stable structured extraction, and avoids the extra runtime and output\-budget complications associated with thinking\-mode models\.
The main goal of this configuration was to make final answering comparable across memory systems\. Each system was allowed to use its own memory mechanism, but probe answering was standardized through the same reader and structured scorer\.
### A\.3Common Evaluation Protocol
All incremental memory systems were evaluated under the same session\-by\-session protocol\. For each episode, we created a fresh isolated memory store, fed the episode sessions in chronological order, retrieved the system’s memory or context at probe time, and then asked the unified reader to answer the probe using that context\.
Formally, for each episode, the evaluation followed:
1. 1\.Reset or create a fresh memory store\.
2. 2\.Ingest sessions sequentially in their original order\.
3. 3\.For each probe, retrieve memory/context from the system rather than giving the full episode\.
4. 4\.Pass the retrieved memory/context to the unified reader\.
5. 5\.Score the structured answer using the same typed\-field scorer\.
Whenever a memory system exposed a context\-only retrieval interface, we used it instead of the system’s own final answer generation\. This ensures that differences in final answers primarily reflect the retrieved or maintained memory rather than differences in the answer generator\. If a system did not expose a clean retrieval interface, we treated it as an end\-to\-end memory system and documented this separately\.
### A\.4Evaluated Systems
We successfully evaluated six incremental memory systems:
- •Mem0, representing extraction\-based long\-term memory\.
- •LangMem, representing framework\-level memory extraction and consolidation\.
- •Graphiti/Zep, representing temporal knowledge\-graph memory\.
- •Cognee, representing graph\-vector hybrid memory\.
- •A\-MEM, representing agentic self\-organizing memory\.
- •MemoryOS, representing hierarchical memory\-OS style memory\.
We also evaluated retrieval\-based diagnostic baselines, including Naive RAG, Time\-aware RAG, and Oracle RAG\. Naive RAG retrieves top\-ranked sessions for each probe\. Time\-aware RAG additionally annotates retrieved sessions with their original temporal position\. Oracle RAG bypasses retrieval and supplies clean non\-filler evidence sessions, serving as a diagnostic upper bound rather than a deployable baseline\.
### A\.5Systems Considered but Excluded
We attempted to evaluate several additional memory systems but excluded them due to implementation or protocol incompatibilities\.
- •MemGPT\([Packer et al\., 2024](https://arxiv.org/html/2609.30558#bib.bib17)\)was excluded because the available implementation required a different version of python/server setup and its tool\-based memory mechanism was incompatible with our Gemini\-controlled function\-calling protocol\.
- •Hindsight\([Latimer et al\., 2025](https://arxiv.org/html/2609.30558#bib.bib34)\)was excluded due to database and system dependency constraints, including PostgreSQL/pgvector and glibc incompatibilities in our environment\.
- •Memory\-R1\([Yan et al\., 2026](https://arxiv.org/html/2609.30558#bib.bib18)\)was excluded because a runnable implementation was not available at the time of our experiments, and the method is an RL fine\-tuning framework rather than an inference\-time memory system compatible with our protocol\.
- •Memoria\([Sarin et al\., 2025](https://arxiv.org/html/2609.30558#bib.bib32)\)was excluded due to dependency conflicts with the unified reader stack and the absence of a clean fixed\-transcript ingestion API under our setting\.
These exclusions do not change the goal of our evaluation, which is to diagnose representative memory mechanisms rather than exhaustively rank all available systems\.MemProbeis intended to diagnose representative memory mechanisms rather than exhaustively rank all available memory products or research prototypes\.
### A\.6Artifacts
For reproducibility, we release the episode suite, latent specifications, runner scripts, raw prediction reports, metric computation scripts, and the full implementation log\. The full implementation log contains exact versions, installation notes, system\-specific patches, runtime characteristics, and detailed failure modes for excluded systems\.
## Appendix BDemonstration Suite Details
Section[3](https://arxiv.org/html/2609.30558#S3)definesMemProbeas a set of diagnostic paradigm specifications\. This appendix describes one concrete instantiation of these specifications: a 56\-episode demonstration suite designed to evaluate stability–plasticity behavior in agent memory systems\. The suite is not intended to exhaust all possible memory scenarios\. Rather, it serves as a controlled and reproducible testbed for showing how theMemProbeparadigms can be instantiated, validated, and used to produce diagnostic behavioral profiles\.
### B\.1Suite Overview
The demonstration suite contains 56 episodes, 1,246 ordered sessions, and 224 probe questions\. Each episode instantiates one of the fourMemProbeparadigms: Interference, Misinformation, Consolidation Strength, and Reconsolidation Window\. Each episode contains an initial target memory, a controlled perturbation condition, naturalistic filler interactions, and probes targeting different behavioral readouts\.
Suite propertyValueNumber of episodes56Number of sessions1,246Number of probes224Paradigms4DomainsPersonal, Work, AgenticSplit\-perturbation episodes12Non\-split episodes44Probe typesCurrent, Previous, Change, Source, Conflict, Temporal
Table 8:Summary statistics of the 56\-episodeMemProbedemonstration suite\.The suite is balanced across the four diagnostic paradigms, with 14 episodes per paradigm\. We also balance update\-oriented and preserve\-oriented conditions so that simple strategies such as always updating or always preserving cannot solve the suite\. In addition, the suite contains both local\-evidence episodes, where the relevant perturbation is contained in a single interaction, and split\-evidence episodes, where the trigger and accepted new value are distributed across multiple sessions\.
The 56\-episode suite is a reference instantiation ofMemProberather than the specification itself\. Other researchers may instantiate the same paradigms with different domains, languages, interaction styles, or memory systems, as long as the controlled relation among the initial memory, perturbation, expected behavior, and probes is preserved\.
### B\.2Domains and Fact Types
To avoid limiting evaluation to lifestyle preferences, we instantiate the paradigms across three broad domains: personal memory, work or productivity memory, and agentic or tool\-state memory\. This reflects the fact that modern agent memory systems are used not only to remember user preferences, but also to maintain project state, configuration values, tool outputs, and task\-specific commitments\.
DomainExample fact typesPersonal / lifestyleCurrent gym, dietary preference, coffee preference, commute plan, routine preferenceWork / productivitySubmission deadline, project owner, team standup time, deployment cadence, review scheduleAgentic / tool\-stateReported port number, active branch, selected database, retry limit, configuration value, tool\-reported stateTable 9:Domains and representative fact types used in the demonstration suite\.Each fact type is chosen to support both update and preserve conditions\. For example, a project owner can change through a confirmed handoff, but a teammate’s temporary involvement should not be treated as an ownership transfer\. Similarly, a reported port number may be updated by an authoritative deployment log, while a stale tool output should not be accepted as the current state\. These distinctions allow similar surface values to play different roles depending on source, status, and target applicability\.
### B\.3Construction Pipeline
The suite was constructed using a staged pipeline that separates experimental design from surface dialogue generation\. This separation is important because the main object of evaluation is not the wording of any single episode, but the controlled relationship between initial memory, perturbation, expected behavior, and probes\.
The construction pipeline consists of five stages:
1. 1\.Fact\-type design\.We manually define fact types compatible with one or moreMemProbeparadigms\. Each fact type must support clear initial values, plausible updates, plausible distractors, and unambiguous current\-state labels\.
2. 2\.Value\-pair design\.For each fact type, we define initial values, candidate new values, distractors, and non\-target values\. Value pairs are reviewed to avoid ambiguity, trivial templates, and obvious shortcuts\.
3. 3\.Latent specification\.Each episode is first represented as a structured latent specification\. The specification fixes the paradigm, condition, expected behavior, canonical values, source type, support level, reactivation status, and required probe types\.
4. 4\.Dialogue generation\.Dialogue sessions are generated from the latent specification\. Generation is constrained so that the target memory, perturbation, fillers, and probes remain consistent with the structured design\. Some filler interactions were adapted from or inspired by public long\-memory benchmark materials, including LongMemEval[Wu et al\. \(2024\)](https://arxiv.org/html/2609.30558#bib.bib1), LoCoMo[Maharana et al\. \(2024\)](https://arxiv.org/html/2609.30558#bib.bib2), and MemoryArena[He et al\. \(2026\)](https://arxiv.org/html/2609.30558#bib.bib5), while the target memory manipulations, perturbations, probes, and gold labels were constructed according to ourMemProbelatent specifications\.
5. 5\.Validation and cleanup\.We apply rule\-based checks, structured gold checks, LLM\-judge consistency checks, Full\-context reader validation, retrieval\-based baselines, and heuristic baselines\. Episodes with schema inconsistencies, ambiguous gold labels, or accidental updates are revised\.
This pipeline keeps the manipulated experimental variable separate from surface dialogue form\. It also makes the suite extensible: new domains or fact types can be added by creating new latent specifications that satisfy the same paradigm constraints\.
### B\.4Latent Specifications
Each episode is grounded in a latent specification that records the controlled variables of the trial\. The latent specification is used both for generation and for evaluation\. It prevents the dialogue from being treated as unstructured text and ensures that each episode has a well\-defined expected behavior\.
Each latent specification contains the following fields:
- •Paradigm and condition: theMemProbeparadigm and the specific condition being instantiated, such asauthoritative\_update,decoy\_no\_update,assistant\_noise,stale\_tool,high\_support\_weak\_challenge, orreactivated\_weak\_evidence\.
- •Expected behavior: whether the system should update the target memory or preserve the existing memory\.
- •Canonical values: the initial value, accepted new value if applicable, conflict value, distractor values, and previous value\.
- •Source and status: whether the perturbation comes from the user, assistant, third party, tool, stale tool output, or another entity\.
- •Support and reactivation variables: for consolidation and reconsolidation trials, the support level, evidence strength, and whether the prior memory is explicitly reactivated\.
- •Split perturbation metadata: whether the perturbation evidence is contained in one session or split across multiple sessions\.
- •Structured gold: canonical answers for current value, previous value, change status, source, conflict status, and temporal sequence\.
### B\.5Dialogue and Perturbation Structure
Each episode consists of ordered sessions\. Sessions are assigned both a global index and a phase label\. The main phases are encoding, filler, perturbation, post\-perturbation context, and probing\.
Encoding sessions establish the initial memory\. Filler sessions introduce natural background interactions and near\-miss facts\. Confusable sessions mention values that are semantically or lexically similar to the target fact but should not determine the gold answer\. Perturbation sessions introduce the critical update, conflict, weak evidence, or reactivation\. Post\-perturbation sessions test whether systems are overly sensitive to recency or surface similarity\.
A standard non\-split episode follows the pattern:
Encoding→Filler→Perturbation\\displaystyle\\text\{Encoding\}\\rightarrow\\text\{Filler\}\\rightarrow\\text\{Perturbation\}\(6\)→Post\-perturbation→Probing\.\\displaystyle\\rightarrow\\text\{Post\-perturbation\}\\rightarrow\\text\{Probing\}\.
In split\-perturbation episodes, the perturbation is intentionally distributed:
Encoding→Filler→P1\\displaystyle\\text\{Encoding\}\\rightarrow\\text\{Filler\}\\rightarrow P\_\{1\}\(7\)→Interleaving filler→P2\\displaystyle\\rightarrow\\text\{Interleaving filler\}\\rightarrow P\_\{2\}→Post\-perturbation→Probing\.\\displaystyle\\rightarrow\\text\{Post\-perturbation\}\\rightarrow\\text\{Probing\}\.
Here,P1P\_\{1\}may contain the source, trigger, or reason for change, whileP2P\_\{2\}contains the accepted current value or final decision\. This design prevents a retrieval system from solving the episode by retrieving a single decisive update sentence\. Instead, the system must integrate evidence across multiple sessions\.
Split perturbation is included to test distributed evidence integration\. In many real interactions, an update is not expressed as a single sentence of the form “the value changed fromxxtoyy\.” A user may first indicate that a plan has changed, then later mention the accepted new value\. Alternatively, a tool may provide a trigger, while a user later confirms how it should be interpreted\.
The split design separates three pieces of information that are often collapsed in simple benchmarks:
- •Change trigger: evidence that some update or reconsideration occurred\.
- •Source or authority: who or what introduced the relevant evidence\.
- •Accepted current value: the value that should be treated as current\.
A system that only retrievesP1P\_\{1\}may know that a change occurred but fail to identify the new value\. A system that only retrievesP2P\_\{2\}may identify the new value but miss the trigger or source\. A robust memory system should connect these pieces of evidence while preserving the previous state\.
### B\.6Structured Gold Labels
Each probe is paired with structured gold labels\. We avoid relying only on free\-form textual answers because many correct predictions differ in wording\. Structured gold fields specify the intended answer at the level of value, source, status, target applicability, and temporal relation\.
For current\-value probes, the gold label specifies the accepted current value\. For previous\-value probes, it specifies the earlier value that should remain historically accessible\. For change\-detection probes, it specifies whether a confirmed change occurred\. For source probes, it specifies who or what introduced the relevant evidence\. For conflict\-value probes, it specifies the conflicting value, its source, its status, and whether it applies to the target fact\. For temporal probes, it specifies the old\-to\-new direction, trigger, reason, and whether an exact time is required\.
For example, a conflict gold label has the following abstract structure:
```
{
"conflict_value": "...",
"mentioned": true,
"source": "...",
"status": "<not_accepted | stale |
other_entity | hypothetical>",
"applies_to_target": true or false,
"correct_current_value": "..."
}
```
Similarly, a temporal gold label has the following structure:
```
{
"change_direction": "<old_to_new |
no_confirmed_change>",
"old_value": "...",
"new_value": "...",
"trigger": "...",
"reason": "...",
"exact_time_required": false
}
```
Structured labels make scoring more robust to paraphrase and allow partial subscores\. For example, a system may identify the conflict value but fail to classify whether it was accepted, stale, hypothetical, or associated with another entity\.
## Appendix CProbe Taxonomy and Scoring
This section describes the probe taxonomy and scoring procedure used inMemProbe\. The main text reports aggregate and behavioral metrics, while the details below specify how individual model responses are matched against structured gold labels\.
### C\.1Probe Types
MemProbeuses targeted probes as behavioral readouts rather than ordinary question\-answering items\. Each probe type isolates a different aspect of memory behavior\.
#### Current\-value probes\.
Current\-value probes ask for the currently valid value of the target fact\. They evaluate whether the system reaches the appropriate final memory state after the episode, either by accepting a valid update or preserving the original value when no valid update occurred\.
#### Previous\-value probes\.
Previous\-value probes ask for the earlier value after a change\. They evaluate whether historical information remains accessible after an update, rather than being erased or overwritten by the current value\.
#### Change\-detection probes\.
Change\-detection probes ask whether a real update occurred\. They distinguish systems that simply output a plausible value from systems that recognize the update status of a memory\.
#### Source probes\.
Source probes ask who or what introduced a relevant value\. They evaluate whether the system preserves provenance information, such as whether a value came from the user, an assistant statement, a tool result, a retrieved document, or another source\.
#### Conflict\-value probes\.
Conflict\-value probes ask whether a competing value was mentioned, where it came from, and whether it applied to the target fact\. They test whether the system distinguishes valid evidence from distractors, stale information, misinformation, or values belonging to another entity\.
#### Temporal probes\.
Temporal probes ask about the update sequence, including the original value, current value, trigger, and reason for change\. They evaluate whether the system can reconstruct the temporal structure behind a memory update, not only the final state\.
Together, these probes separate memory behaviors that would otherwise be collapsed by final\-answer accuracy\. A system may recover the current value but lose the previous value, detect that a conflicting value was mentioned but misattribute its source, or identify a change without correctly explaining why the change occurred\.
### C\.2Scoring Procedure
We score six probe types: current value, previous value, change detection, source, conflict value, and temporal sequence\.
#### Current\-value scoring\.
A current\-value response is correct if it identifies the accepted current state of the target fact according to the episode gold label\.
#### Previous\-value scoring\.
A previous\-value response is correct if it recovers the earlier value after an update\. If no update occurred and the probe asks for a previous value, the expected answer is determined by the episode label, such as “no previous value” or the original value when the wording asks for the initially established state\.
#### Change\-detection scoring\.
A change\-detection response is correct if it correctly determines whether a confirmed update occurred\. This score evaluates update\-status recognition rather than value extraction alone\.
#### Source scoring\.
A source response is correct if it identifies the source associated with the relevant value or update\. Acceptable answers may include normalized source categories such as user, assistant, tool, document, third party, stale record, or other entity, depending on the episode label\.
#### Conflict\-value scoring\.
Conflict\-value probes are decomposed into multiple fields: value identification, source identification, status classification, and target applicability\. A response may receive partial credit if it identifies the competing value but misclassifies whether that value was accepted, rejected, stale, hypothetical, unreliable, or non\-applicable to the target entity\.
#### Temporal scoring\.
Temporal probes are scored using the old value, current value, update direction, trigger, and reason for change\. Exact dates or timestamps are required only when explicitly specified by the gold label\. Otherwise, a correct response must capture the update sequence and the event or source that triggered the change\.
### C\.3Normalization and Partial Credit
Before scoring, we normalize common surface variants that do not change meaning\. For example, optional prepositions in time expressions may be ignored, and numeric values may be compared with or without common units when the unit is unambiguous\. For sentence\-level predictions, we extract candidate values when possible before comparison\.
Partial credit is used only for structured probes with multiple fields, such as conflict\-value and temporal probes\. This prevents formatting differences from being counted as semantic errors while preserving important distinctions among value identification, source attribution, status classification, and temporal reasoning\.
### C\.4Execution Errors and Exclusions
API failures, parsing failures, or execution errors unrelated to episode content are marked separately and excluded from the corresponding accuracy denominators\. These exclusions are not treated as model reasoning errors\. All exclusions are reported in the validation tables associated with the experiment\.
## Appendix DInter\-Axis Correlation
MemProbereports results along two decomposition axes: the four cognitive paradigms and the six probe types\. A natural question is how much independent signal each axis carries\. Treating the six memory systems as samples, we compute pairwise correlations between axes and summarize the off\-diagonal entries\.
PearsonSpearmanAxismeanminmeanmin4 paradigms0\.930\.850\.880\.776 probe types0\.780\.430\.750\.43Table 10:Off\-diagonal correlation summary across the two decomposition axes, computed over the six incremental memory systems\.The paradigm axis is highly collinear: all twelve off\-diagonal correlations exceed0\.850\.85, so a system that is strong on one paradigm tends to be strong on all four\. Paradigm\-level accuracy therefore largely reflects general competence rather than four separable capabilities\. This is expected, and it is consistent with the role the paradigms play inMemProbe: they construct mechanistically distinct controlled conditions under which memory behavior can be observed, not four independent scoring dimensions\.
The probe\-type axis is measurably less collinear\. Its minimum off\-diagonal correlation \(0\.430\.43\) is well below the paradigm minimum \(0\.850\.85\), indicating that some probe types capture behavior the others do not predict\. Temporal fidelity is the clearest case: it correlates onlyr=0\.64r=0\.64with current\-value accuracy, so reconstructing the update sequence is substantially less correlated to reporting the correct current value\. Source fidelity, by contrast, correlates more strongly with current value \(r=0\.89r=0\.89\), but this reflects a uniform deficit rather than a shared capability: every system scores far lower on source than on current value \(Table[5](https://arxiv.org/html/2609.30558#S5.T5)\)\. The independent diagnostic signal inMemProbetherefore comes primarily from decomposing correctness by probe type\.
## Appendix ESuite Quality Control and Validation
We validate the suite at three levels: schema consistency, gold\-label consistency, and behavioral sanity checks\. The goal is to ensure that the episodes are internally consistent, answerable under complete evidence, and not solvable by trivial shortcuts\.
### E\.1Schema and Gold\-Label Validation
All episodes are checked for required top\-level fields, latent\-spec fields, session fields, and metadata fields\. Each session has a session index, phase label, session type, dialogue content, and metadata\. Split episodes explicitly record the number of perturbation sessions and the identities of perturbation parts\.
We use structured gold labels for all probe types and manually audit cases where gold labels and dialogue content disagree\. We also normalize common value variants, such as optional prepositions in time expressions and unit\-bearing numeric values\. This reduces false negatives caused by surface mismatch rather than model error\.
### E\.2Probe\-Level Validation
Table[11](https://arxiv.org/html/2609.30558#A5.T11)reports validation accuracy by probe type\. High judge and Full\-context scores indicate that most gold labels are recoverable under complete evidence\. The lower RAG scores on current\-value and change\-detection probes suggest that the suite selectively stresses retrieval and evidence integration rather than merely increasing surface ambiguity\.
Probe typeJudgeFull\-contextNaive RAGCurrent value98\.2%96\.4%71\.4%Previous value100\.0%100\.0%92\.9%Change detection100\.0%94\.6%65\.5%Temporal86\.4%85\.7%81\.8%Source97\.1%97\.1%97\.1%Conflict value100\.0%96\.4%100\.0%
Table 11:Probe\-level validation accuracy for the 56\-episode suite\. Judge and Full\-context scores are high for most probe types, indicating that structured gold labels are largely recoverable when all evidence is visible\. Naive RAG shows the largest degradation on current\-value and change\-detection probes, suggesting that retrieval\-based systems struggle most when they must identify the accepted current state and determine whether a real update occurred\.
### E\.3Behavioral Sanity Checks
We run several reference baselines before using the suite to evaluate memory systems\. A Full\-context reader checks whether the episode is solvable when all evidence is visible\. A naive RAG baseline checks whether retrieval and evidence integration are non\-trivial\. Heuristic baselines such as latest\-value, always\-update, and always\-preserve test whether the suite can be solved by simple shortcuts\.
Table[12](https://arxiv.org/html/2609.30558#A5.T12)summarizes the current validation results\. The judge and Full\-context reader indicate that the episodes are largely readable and internally consistent\. The RAG gap and low heuristic scores indicate that the suite is not solved by simple retrieval or recency shortcuts\.
BaselineAccuracyLLM judge97\.8%Full\-context reader95\.5%Naive RAG81\.2%Latest\-value heuristic18\.8%FC–RAG gap14\.3 ppTable 12:Validation results for the 56\-episode demonstration suite\. The judge and Full\-context reader indicate that the episodes are largely readable and internally consistent, while the RAG gap and low heuristic score indicate that the suite is not solved by simple retrieval or recency shortcuts\.
### E\.4Split\-Perturbation Validation
The split\-perturbation subset provides an additional diagnostic check\. Full\-context performance remains high on split episodes, indicating that these episodes are still answerable when all evidence is visible\. However, naive RAG drops substantially on split episodes, showing that retrieval systems struggle when evidence is distributed across sessions\.
SubsetJudgeFull\-contextNaive RAGSplit episodes95\.8%93\.6%56\.2%Non\-split episodes98\.3%96\.0%88\.0%
Table 13:Split versus non\-split validation results\. Split perturbations preserve Full\-context readability but sharply reduce RAG performance, indicating that they diagnose distributed evidence integration rather than simply making the text ambiguous\.
### E\.5Cross\-model validation
To check whether the validation results depend on a single reader model, we repeat the suite validation with three readers: GPT\-5\.4, Gemini\-3\-Flash, and Qwen3\-32B\. All models are evaluated with the same structured JSON scoring protocol\. The probes remain natural\-language questions, but model responses are normalized into structured fields before scoring\. This reduces surface\-form bias across models while preserving the semantic requirements of each probe\.
Table[14](https://arxiv.org/html/2609.30558#A5.T14)summarizes the overall results\. All three readers achieve high judge accuracy, indicating that the structured gold labels are largely recoverable across models\. Full\-context performance is consistently higher than naive RAG, showing that retrieval\-based reading introduces a substantial additional bottleneck\. Qwen3\-32B is weaker overall, especially on source and conflict\-status reasoning, but it preserves the same qualitative pattern: Full\-context reading outperforms naive RAG\.
BaselineGPT\-5\.4Gemini\-3\-FlashQwen3\-32BJudge96\.994\.290\.2Full\-context92\.488\.878\.1Naive RAG68\.868\.359\.4FC–RAG gap23\.620\.518\.7
Table 14:Cross\-model validation summary under structured JSON scoring\. All three readers show the same qualitative trend: Full\-context performance is substantially higher than naive RAG, indicating that the suite exposes retrieval and evidence\-integration difficulty beyond Full\-context readability\.Table[15](https://arxiv.org/html/2609.30558#A5.T15)breaks down Full\-context performance by probe type\. Across models, current\-value, previous\-value, and change\-detection probes are consistently strong, showing that the core factual memory state is readable under full evidence\. The largest cross\-model differences appear in source and conflict\-value probes\. Qwen3\-32B performs competitively on state\-oriented probes but drops sharply on provenance and conflict\-status probes, suggesting that source attribution and target\-applicability classification are harder than factual value extraction for weaker open\-weight readers\.
Probe typeGPT\-5\.4Gemini\-3\-FlashQwen3\-32BCurrent value96\.492\.994\.6Previous value100\.0100\.096\.4Change detection96\.498\.294\.6Source76\.573\.541\.2Conflict value92\.978\.642\.9Temporal86\.477\.372\.7Overall92\.488\.878\.1
Table 15:Full\-context accuracy by probe type across reader models\. State\-oriented probes are robust across models, whereas source and conflict\-value probes reveal larger differences, especially for Qwen3\-32B\. This suggests that provenance and conflict\-status reasoning are more demanding than extracting current or previous values\.Overall, the cross\-model results support two conclusions\. First, the suite is not specific to the GPT\-5\.4 reader: Gemini\-3\-Flash and Qwen3\-32B recover the main factual memory states under full context, and Gemini closely matches GPT\-5\.4 in the overall validation pattern\. Second, the split\-perturbation effect is robust, retrieval\-based reading fails much more severely on split episodes than on non\-split episodes, especially for probes that require identifying the accepted current value or detecting that a true update occurred\. This validates split perturbation as a diagnostic stressor for distributed evidence integration\.
## Appendix FA Second Suite Instantiation
MemProbeis specified as a construction protocol rather than a fixed dataset\. To verify that the protocol is portable rather than tailored to the 56\-episode suite, we instantiated a second and independent suite using the same five\-stage pipeline \(App\.[B](https://arxiv.org/html/2609.30558#A2)\) and evaluated the same six memory systems on it\.
### F\.1Construction
The second suite contains 40 episodes, 10 per paradigm, balanced across the same three domains\. Two properties make it an independent test of the protocol rather than a re\-wording of the original suite\. First, it uses 12 new fact types with no overlap with the original suite, extending the work and agentic domains in particular\. Second, dialogue surface form was generated with a different generator model \(Gemini\-3\.1\-Pro\), so the instantiation does not inherit the wording conventions of the original suite\. As before, all controlled variables \(latent specifications, perturbations, expected behaviors and gold labels\) were human\-designed and audited rather than generated\.
### F\.2Results
Because the second suite was generated with a Gemini\-family model, we report results under a GPT\-5\.6 reader, keeping generation and reading in different model families as in the main evaluation\.
SystemInterf\.Misinf\.Consol\.Recons\.Overall \[95% CI\]A\-MEM97\.589\.597\.397\.595\.5 \[91\.7, 98\.7\]Cognee86\.592\.192\.594\.691\.4 \[86\.3, 96\.1\]MemoryOS84\.676\.968\.480\.077\.6 \[70\.4, 84\.3\]LangMem82\.570\.386\.863\.275\.8 \[69\.4, 82\.6\]Graphiti75\.071\.166\.783\.874\.0 \[66\.4, 81\.8\]Mem045\.034\.243\.248\.642\.8 \[36\.7, 48\.4\]
Table 16:Paradigm\-wise and overall accuracy on the second suite\. Confidence intervals are episode\-level bootstrap intervals over the 40 episodes\.The core diagnostic conclusions replicate\. A\-MEM and Mem0 again occupy the top and bottom of the ranking, and the source accuracy again falls well below current\-value accuracy for every system\. Absolute scores are higher than on the original suite, which is consistent with the second suite having shorter episodes and a stronger reader\.
One difference is that MemoryOS ranks third here but fifth on the original suite\. The second suite is weighted toward agentic, shorter\-horizon fact types, which suits MemoryOS’s hierarchical short/mid/long\-term organization, whereas the original suite contains longer episodes with more competing\-fact interference\.
## Appendix GLimitations of the Demonstration Suite
The demonstration suite is LLM\-assisted and intentionally controlled\. This improves interpretability and validation, but does not capture the full diversity of real user logs\. The covered domains are representative rather than exhaustive, and the suite does not include all possible uses of long\-term agent memory\.
The current instantiation is also text\-only and English\-only\. It does not evaluate multimodal memory, multilingual interaction, or memory updates grounded in long\-running external environments\.
These limitations motivate larger and noisier instantiations rather than changing the role ofMemProbe\. Future work can instantiate the same paradigm specifications in multilingual, multimodal, user\-derived, or application\-specific settings while preserving the controlled variables that make stability–plasticity behavior measurable\.
## Appendix HLicense and Artifact Usage
### H\.1License
We use standard licenses from the community and provide the following links to the licenses for the datasets, codes, and models that we used in this paper:
LongMemEval:[MIT](https://github.com/xiaowu0162/LongMemEval/blob/main/LICENSE)
### H\.2Artifact Usage
The use of existing artifacts is consistent with their intended use in this work\. We will make our code and models publicly accessible and all created artifacts will be only for research purposes and should not be used outside of research contexts\.
### H\.3AI Assistants Usage
We strictly follow the ARR rules in AI Assistants usage and carefully check all the produced artifacts\.Similar Articles
MEMPROBE: Probing Long-Term Agent Memory via Hidden User-State Recovery
MEMPROBE is a benchmark that evaluates long-term memory in LLM agents by reconstructing hidden user states from the agent's memory after interaction.
Self-Improving Memory for Agents (6 minute read)
Perplexity Brain is a memory system that builds a persistent context graph across tasks, projects, decisions, files, and sources, enabling agents to start with relevant context instead of from scratch, improving answer correctness and reducing task costs.
Grounding Agent Memory: Environment-Probing Curation for Enterprise Agents
This paper introduces environment-probing curation to improve persistent memory for enterprise agents, showing substantial gains in task performance and cost reduction on benchmarks like CLBench and APEX.
Personalize-then-Store: Benchmarking and Learning Personalized Memory for Long-horizon Agents
This paper introduces PerMemBench, the first benchmark for evaluating personalized memory systems in LLM-based agents, and proposes a session-level storage gating framework that adapts memory policies to individual user contexts.
@omarsar0: // AutoMem // I quite like this idea of metamemory. (bookmark it) This new research from Stanford treats agent's memory…
This Stanford research paper introduces AutoMem, a framework that treats agent memory management as a trainable skill. By optimizing memory structure and proficiency separately, AutoMem improves base agent performance 2x-4x on long-horizon tasks, enabling a 32B open-weight model to compete with frontier systems like Claude Opus 4.5 and Gemini 3.1 Pro Thinking.