CueMem: Cue-Guided Context Reconstruction for Long-Term Conversational Memory

arXiv cs.CL Papers

Summary

CueMem is a cue-guided framework for long-term conversational memory that reconstructs dialogue context using retrieval cues to improve performance over existing baselines in LLM-powered agents.

arXiv:2609.12354v1 Announce Type: new Abstract: Long-term conversational agents must answer user queries by recalling information from extended dialogue histories, yet directly using the full history is costly and often unreliable, while compressed memory units may lose fine-grained evidence needed for question answering. Motivated by the reconstructive view of autobiographical memory, we propose CueMem, a cue-guided framework that treats extracted memory records as retrieval cues rather than self-contained evidence and reconstructs query-relevant dialogue context from their source turns. During memory construction, CueMem extracts fine-grained memory cues from dialogue turns and links each cue to its source turn. At query time, it retrieves query-relevant cues, maps them to source-turn anchors, and expands from these anchors over a turn graph that captures temporal proximity and semantic relatedness, reconstructing a compact evidence context from the original dialogue for LLM answer generation. Experiments on LoCoMo and LongMemEval show that CueMem consistently outperforms representative long-term memory baselines. Further analyses show that graph-based context reconstruction helps recover supporting dialogue evidence while reducing query-time input tokens and latency compared with the full-history LLM setting. These results highlight retrieval cues as an effective alternative to self-contained memory evidence for long-term conversational question answering.
Original Article
View Cached Full Text

Cached at: 09/14/26, 08:34 AM

# CueMem: Cue-Guided Context Reconstruction for Long-Term Conversational Memory
Source: [https://arxiv.org/html/2609.12354](https://arxiv.org/html/2609.12354)
Changjian WangAffiliation:Mashang Consumer Finance Co\., Ltd\.Affiliation:Harbin Institute of Technology, ShenzhenEmail:[changjian\.wang@msxf\.com](mailto:[email protected])Weili GuanAffiliation:Harbin Institute of Technology, ShenzhenShuming ShiAffiliation:Mashang Consumer Finance Co\., Ltd\.Quan LuAffiliation:Mashang Consumer Finance Co\., Ltd\.Ning JiangAffiliation:Mashang Consumer Finance Co\., Ltd\.

###### Abstract

Long\-term conversational agents must answer user queries by recalling information from extended dialogue histories, yet directly using the full history is costly and often unreliable, while compressed memory units may lose fine\-grained evidence needed for question answering\. Motivated by the reconstructive view of autobiographical memory, we propose CueMem, a cue\-guided framework that treats extracted memory records as retrieval cues rather than self\-contained evidence and reconstructs query\-relevant dialogue context from their source turns\. During memory construction, CueMem extracts fine\-grained memory cues from dialogue turns and links each cue to its source turn\. At query time, it retrieves query\-relevant cues, maps them to source\-turn anchors, and expands from these anchors over a turn graph that captures temporal proximity and semantic relatedness, reconstructing a compact evidence context from the original dialogue for LLM answer generation\. Experiments on LoCoMo and LongMemEval show that CueMem consistently outperforms representative long\-term memory baselines\. Further analyses show that graph\-based context reconstruction helps recover supporting dialogue evidence while reducing query\-time input tokens and latency compared with the full\-history LLM setting\. These results highlight retrieval cues as an effective alternative to self\-contained memory evidence for long\-term conversational question answering\.

## 1Introduction

Figure 1:An example comparison between memory units and memory cues\. A coarse memory unit may lose fine\-grained semantic information due to compression or mixed semantic representations, e\.g\., a “travel planning” memory unit may merge destination, budget, and hotel preferences, making it difficult to obtain the evidence for the user’s query\. In contrast, fine\-grained memory cues linked to source turns provide focused retrieval targets, enabling the system to recover the relevant source turn T2 and reconstruct the dialogue context needed to answerLakeside Town\.Large Language Models \(LLMs\) have developed rapidly in recent years\([Achiam et al\., 2023](https://arxiv.org/html/2609.12354#bib.bib19);[Yang et al\., 2025](https://arxiv.org/html/2609.12354#bib.bib20)\), leading to substantial improvements in the capabilities of LLM\-powered conversational agents\. In real\-world scenarios, conversations are often long\-running and multi\-turn, requiring agents to answer user queries by effectively leveraging long\-term dialogue history\. Naively incorporating the full interaction history is often ineffective due to context length limitations, noisy information, the “lost\-in\-the\-middle” phenomenon\([Liu et al\., 2024](https://arxiv.org/html/2609.12354#bib.bib11)\), and increased inference latency\. Therefore, how to model and utilize long\-term conversational memory has become a major challenge for conversational agents\.

Existing methods typically construct retrievable memory units from dialogue history through segmentation, compression, summarization, or salient information extraction, and retrieve the most relevant units as context for generation\([Zhong et al\., 2024](https://arxiv.org/html/2609.12354#bib.bib17);[Pan et al\., 2025](https://arxiv.org/html/2609.12354#bib.bib21);[Fang et al\., 2026](https://arxiv.org/html/2609.12354#bib.bib5)\)\. However, such memory units often remain compressed or decontextualized fragments of the original dialogue, and vector\-based indexing may entangle multiple semantic signals within a single dense representation\. More importantly, many existing systems treat retrieved memory units as self\-contained evidence for generation\. This design contrasts with the reconstructive view of autobiographical memory, which suggests that remembering is guided by cues and reconstructed from broader autobiographical knowledge stores\([Conway and Rubin, 1993](https://arxiv.org/html/2609.12354#bib.bib18);[Conway and Pleydell\-Pearce, 2000](https://arxiv.org/html/2609.12354#bib.bib1);[Holland and Kensinger, 2010](https://arxiv.org/html/2609.12354#bib.bib2)\)\.

For example, as shown in Figure[1](https://arxiv.org/html/2609.12354#S1.F1), an existing memory system may group several turns about destination, budget, and hotel preference into a single memory unit about “travel planning”\. However, a later query such as “Which place we mentioned earlier would my mother prefer?” targets only a fine\-grained aspect of this topic, making retrieval challenging because the relevant signal is mixed with multiple other semantics within the same retrievable unit\. In contrast, fine\-grained cues such as “the user’s mother prefers calm places” provide more focused retrieval targets and can guide the system back to the original dialogue turns for context reconstruction\. Inspired by this cognitive perspective, we argue that long\-term conversational memory should follow a similar retrieval\-and\-reconstruction process: first identifying fine\-grained cues that match the current query, and then recovering the original dialogue context around these cues before generating the answer\.

In this paper, we propose CueMem, a simple and effective long\-term memory system for agent conversations, inspired by the reconstructive view of autobiographical memory\. CueMem first constructs fine\-grained memory cues from dialogue turns, with each cue serving as a semantically focused retrieval target linked to its source turn\. Given a user query, CueMem retrieves the most relevant cues and uses their source turns as anchors for context reconstruction\. To recover supporting evidence, CueMem builds a turn graph that connects dialogue turns based on temporal proximity and semantic similarity, and expands from the anchored turns over this graph to collect query\-relevant dialogue context\. Finally, the reconstructed context is provided to the LLM as memory evidence for answer generation\. We evaluate CueMem on two long\-term conversational question answering benchmarks, LoCoMo\([Maharana et al\., 2024](https://arxiv.org/html/2609.12354#bib.bib3)\)and LongMemEval\([Wu et al\., 2025](https://arxiv.org/html/2609.12354#bib.bib4)\), and show that it consistently outperforms baseline memory models\. These results demonstrate the effectiveness of retrieving fine\-grained memory cues and reconstructing dialogue context from anchored turns\.

Our contributions are summarized as follows:

- •We introduce a cue\-centered view of long\-term conversational memory, in which fine\-grained memory cues guide the reconstruction of relevant dialogue context rather than serving as self\-contained evidence for generation\.
- •We propose CueMem, a simple yet effective framework that constructs fine\-grained memory cues from dialogue turns, retrieves query\-relevant cues, and recovers supporting dialogue context by expanding from their source\-turn anchors on a turn\-level graph\.
- •Experiments on two long\-term conversational question answering benchmarks show that CueMem consistently outperforms strong long\-term memory baselines, demonstrating the effectiveness of fine\-grained cue retrieval and anchor\-based context reconstruction\.

## 2Related Work

Retrieval\-Augmented Generation \(RAG\) augments language models by retrieving relevant external context before generation\([Lewis et al\., 2020](https://arxiv.org/html/2609.12354#bib.bib9);[Gao et al\., 2023](https://arxiv.org/html/2609.12354#bib.bib10)\)\. In long\-context and conversational settings, RAG is commonly used to retrieve dialogue turns, chunks, summaries, or memory entries, reducing the need to process the full history and alleviating long\-context issues such as the “lost\-in\-the\-middle” problem\([Liu et al\., 2024](https://arxiv.org/html/2609.12354#bib.bib11)\)\. Recent work also explores more structured retrieval mechanisms, such as graph\-based RAG\([Edge et al\., 2024](https://arxiv.org/html/2609.12354#bib.bib12);[Guo et al\., 2024](https://arxiv.org/html/2609.12354#bib.bib13);[Gutiérrez et al\., 2024](https://arxiv.org/html/2609.12354#bib.bib22);[Gutiérrez et al\., 2025](https://arxiv.org/html/2609.12354#bib.bib14)\), which organizes documents or entities into graphs to support multi\-hop or global information access\. However, these RAG systems mainly target static external knowledge bases, and their fixed chunking and retrieval can suffer from granularity mismatch in dynamic long\-term conversations\.

Long\-term conversational memory aims to help LLM agents retain and use information from extended interactions\. Early agent memory systems store observations or dialogue histories as memory streams and retrieve relevant records for future reasoning or response generation\([Park et al\., 2023](https://arxiv.org/html/2609.12354#bib.bib15);[Zhong et al\., 2024](https://arxiv.org/html/2609.12354#bib.bib17);[Packer et al\., 2023](https://arxiv.org/html/2609.12354#bib.bib16)\)\. More recent methods further organize memories through hierarchical modules, structured notes, graph links, or consolidation mechanisms to support scalable storage, updating, and retrieval\([Kang et al\., 2025](https://arxiv.org/html/2609.12354#bib.bib6);[Chhikara et al\., 2025](https://arxiv.org/html/2609.12354#bib.bib7);[Xu et al\., 2025](https://arxiv.org/html/2609.12354#bib.bib8);[Fang et al\., 2026](https://arxiv.org/html/2609.12354#bib.bib5)\)\. However, these methods often rely on compressed summaries or memory notes as the direct evidence for generation, which may lose fine\-grained details and local dialogue context\. Moreover, merging multiple facts into one memory unit can entangle different semantic signals and reduce retrieval precision\.

Cognitive studies of autobiographical memory provide a useful perspective for addressing this issue\. Autobiographical memory is commonly defined as “memory for the events of one’s life”\([Conway and Rubin, 1993](https://arxiv.org/html/2609.12354#bib.bib18)\)\. Although agent memory is not human memory, long\-term conversational memory plays a functionally analogous role: it records events from an agent’s interaction history and supports later recall in response to user queries\. Cognitive studies suggest that autobiographical memories are not stored as perfect records, but are reconstructed from broader autobiographical knowledge with the help of cues\([Holland and Kensinger, 2010](https://arxiv.org/html/2609.12354#bib.bib2);[Conway and Pleydell\-Pearce, 2000](https://arxiv.org/html/2609.12354#bib.bib1)\)\. This perspective motivates CueMem, which treats stored memories as fine\-grained retrieval cues for reconstructing evidence rather than as self\-contained evidence\.

## 3Problem Formulation

We focus on long\-term conversational question answering, where an agent answers a user query given a long dialogue history\. Formally, let𝒟=\{u1,u2,…,uT\}\\mathcal\{D\}=\\\{u\_\{1\},u\_\{2\},\\ldots,u\_\{T\}\\\}denote a dialogue history consisting ofTTturns\. The dialogue history may span multiple sessions, where each session is a temporally contiguous block of turns, but we represent the entire history as a chronological sequence of turns\. Each turnuiu\_\{i\}consists of a user utterance and the corresponding agent response, and may include associated metadata such as timestamp\. Given a user queryqq, the goal is to generate an answeraathat is both relevant toqqand faithful to the information contained in𝒟\\mathcal\{D\}\. Since𝒟\\mathcal\{D\}can be long, noisy, and only partially relevant to the current query, the system should identify a compact evidence context rather than rely on the entire dialogue history as input\. We formulate the task as selecting a query\-relevant evidence contextEq⊆𝒟\\mathrm\{E\}\_\{q\}\\subseteq\\mathcal\{D\}and generating the answer conditioned on both the query and this context:

a=LLM⁡\(Prompt⁡\(q,Eq\)\),a=\\mathrm\{LLM\}\(\\mathrm\{Prompt\}\(q,\\mathrm\{E\}\_\{q\}\)\),whereEq\\mathrm\{E\}\_\{q\}contains the historical information necessary for answeringqq, andPrompt⁡\(⋅\)\\mathrm\{Prompt\}\(\\cdot\)denotes the answer\-generation prompt that formats the query and evidence context for the LLM\.

Figure 2:Overview of CueMem\. In the offline stage, CueMem extracts fine\-grained memory cues from dialogue turns, links each cue to its source turn, encodes the cues into dense vectors for retrieval, and builds a turn graph with temporal and semantic edges\. In the online stage, CueMem encodes the user query into the same vector space, retrieves query\-relevant cues, maps them to source\-turn anchors, expands over the turn graph to reconstruct supporting dialogue context, and feeds the reconstructed context to the LLM for answer generation\.
## 4Methodology

CueMem follows a cue\-to\-anchor\-to\-context memory reconstruction pipeline, as illustrated in Figure[2](https://arxiv.org/html/2609.12354#S3.F2)\. Given a long dialogue history, CueMem first extracts fine\-grained memory cues from individual dialogue turns and links each cue to its source turn\. It then builds a turn\-level graph over the original dialogue history, where temporal edges preserve local conversational continuity and semantic edges connect turns based on cue\-level semantic similarity\. At query time, CueMem retrieves query\-relevant cues, maps them to source\-turn anchors, and expands from these anchors on the turn graph to reconstruct the evidence context\. The reconstructed context is then provided to the LLM for answer generation\.

### 4\.1Memory Cue Extraction

The first step of CueMem is to convert the dialogue history into a set of fine\-grained memory cues\. Unlike segment\- or summary\-level memories that compress multiple semantic signals into a single unit, each cue is designed to capture a relatively atomic piece of information from the dialogue\. Memory cues can take various forms, such as natural language statements or structured relational records\. For simplicity and interpretability, we instantiate each cue as a relational triple in this work\. Formally, given the dialogue history𝒟\\mathcal\{D\}, we extract a cue set𝒞=\{c1,c2,…,cN\}\\mathcal\{C\}=\\\{c\_\{1\},c\_\{2\},\\ldots,c\_\{N\}\\\}\. Each cue is represented asci=\(ti,τi\)c\_\{i\}=\(t\_\{i\},\\tau\_\{i\}\), wheretit\_\{i\}denotes a triple\(subject,relation,object\)\(\\textit\{subject\},\\textit\{relation\},\\textit\{object\}\)extracted from the dialogue andτi∈\{1,2,…,T\}\\tau\_\{i\}\\in\\\{1,2,\\ldots,T\\\}denotes a pointer to the source turn in𝒟\\mathcal\{D\}\. The triple provides a semantically focused representation for retrieval, while the pointer links the cue back to the original dialogue turn for later context reconstruction\. In practice, we use an LLM to extract relational triples from each dialogue turn to construct memory cues\. A turn may yield multiple cues if it contains several useful pieces of information, or no cue if it does not contain any\. We further encode each tripletit\_\{i\}with an embedding model to obtain its dense representation𝐭i=Enc⁡\(ti\)\\mathbf\{t\}\_\{i\}=\\mathrm\{Enc\}\(t\_\{i\}\), which is used for subsequent retrieval and graph construction\.

### 4\.2Turn Graph Construction

After extracting memory cues, CueMem builds a turn\-level graph over the original dialogue history to support later context reconstruction\. The graph preserves dialogue structure beyond isolated cue matches, allowing the system to recover temporally and semantically related turns once a source turn is selected as an anchor\. Formally, we construct a graph𝒢=\(𝒱,ℰ\)\\mathcal\{G\}=\(\\mathcal\{V\},\\mathcal\{E\}\), where each nodevi∈𝒱v\_\{i\}\\in\\mathcal\{V\}corresponds to a dialogue turnui∈𝒟u\_\{i\}\\in\\mathcal\{D\}\. The edge setℰ\\mathcal\{E\}consists of two types of edges: temporal edgesℰtemp\\mathcal\{E\}\_\{\\mathrm\{temp\}\}and semantic edgesℰsem\\mathcal\{E\}\_\{\\mathrm\{sem\}\}\. Temporal edges capture chronological continuity between dialogue turns, while semantic edges capture semantic associations between turns\. We describe the construction of these two edge types below\.

#### Temporal edges\.

Temporal edges capture temporal proximity between dialogue turns\. For each turn nodeviv\_\{i\}, we define its temporal neighbors as the turns within a fixed chronological window:

𝒩temp​\(vi\)=\{vj∈𝒱∣0<\|i−j\|≤w\},\\mathcal\{N\}\_\{\\mathrm\{temp\}\}\(v\_\{i\}\)=\\\{v\_\{j\}\\in\\mathcal\{V\}\\mid 0<\|i\-j\|\\leq w\\\},wherewwis the window size\. For eachvj∈𝒩temp​\(vi\)v\_\{j\}\\in\\mathcal\{N\}\_\{\\mathrm\{temp\}\}\(v\_\{i\}\), we add a directed temporal edge:

\(vi,vj\)∈ℰtemp\.\(v\_\{i\},v\_\{j\}\)\\in\\mathcal\{E\}\_\{\\mathrm\{temp\}\}\.These edges help recover local context around an anchored turn, including turns that may be difficult to retrieve directly but are useful for the LLM to understand the conversational context and generate a faithful answer\.

#### Semantic edges\.

Semantic edges are used to capture non\-local relations between turns that are semantically related but may be far apart in the dialogue history\. We construct these semantic edges by computing similarity at the cue level and then mapping similar cues back to their source turns\. Compared with turn\-level similarity, cue\-level similarity relies on more atomic semantic representations, reducing the interference of mixed semantics and making the resulting associations more focused and precise\. More specifically, for each cueci=\(ti,τi\)c\_\{i\}=\(t\_\{i\},\\tau\_\{i\}\), we retrieve its top\-kknearest cues according to cosine similarity between cue embeddings:

𝒩sem\(ci\)=TopKcj∈𝒞∖\{ci\}cos\(𝐭i,𝐭j\)\.\\mathcal\{N\}\_\{\\mathrm\{sem\}\}\(c\_\{i\}\)=\\mathrm\{TopK\}\_\{c\_\{j\}\\in\\mathcal\{C\}\\setminus\\\{c\_\{i\}\\\}\}\\cos\(\\mathbf\{t\}\_\{i\},\\mathbf\{t\}\_\{j\}\)\.For each retrieved cuecj=\(tj,τj\)∈𝒩sem​\(ci\)c\_\{j\}=\(t\_\{j\},\\tau\_\{j\}\)\\in\\mathcal\{N\}\_\{\\mathrm\{sem\}\}\(c\_\{i\}\), we map both cues to their source\-turn nodes and add a directed semantic edge fromvτiv\_\{\\tau\_\{i\}\}tovτjv\_\{\\tau\_\{j\}\}in the turn graph:

\(vτi,vτj\)∈ℰsem\.\(v\_\{\\tau\_\{i\}\},v\_\{\\tau\_\{j\}\}\)\\in\\mathcal\{E\}\_\{\\mathrm\{sem\}\}\.

### 4\.3Context Reconstruction

Given a user query, CueMem reconstructs the evidence context through a cue\-to\-anchor\-to\-context process\. It first retrieves fine\-grained memory cues that are semantically relevant to the query\. The source turns linked to these retrieved cues are then treated as anchors in the turn graph\. Starting from these anchors, CueMem expands along temporal and semantic edges to recover both local conversational context and non\-local semantically related evidence\. The selected turns are finally assembled as the evidence context for answer generation\.

#### Cue retrieval\.

We first encode the user queryqqwith the same embedding model used for memory cues, obtaining𝐪=Enc⁡\(q\)\\mathbf\{q\}=\\mathrm\{Enc\}\(q\)\. CueMem then retrieves the top\-mmcues that are most similar to the query:

𝒞q=TopKci∈𝒞cos\(𝐪,𝐭i\),\\mathcal\{C\}\_\{q\}=\\mathrm\{TopK\}\_\{c\_\{i\}\\in\\mathcal\{C\}\}\\cos\(\\mathbf\{q\},\\mathbf\{t\}\_\{i\}\),wheremmis the number of retrieved cues\. The source\-turn pointers of these cues define the anchor set:

𝒜q=\{τi∣ci∈𝒞q\}\.\\mathcal\{A\}\_\{q\}=\\\{\\tau\_\{i\}\\mid c\_\{i\}\\in\\mathcal\{C\}\_\{q\}\\\}\.

#### Graph expansion\.

Starting from the anchor set𝒜q\\mathcal\{A\}\_\{q\}, CueMem expands over the turn graph𝒢\\mathcal\{G\}to collect supporting turns connected by temporal and semantic edges\. We first collect the one\-hop graph neighbors of the anchors:

𝒩𝒢\(𝒜q\)=\{j∣∃i∈𝒜q,\(vi,vj\)∈ℰtemp∪ℰsem\}\.\\mathcal\{N\}\_\{\\mathcal\{G\}\}\(\\mathcal\{A\}\_\{q\}\)=\\\{j\\mid\\exists i\\in\\mathcal\{A\}\_\{q\},\\ \(v\_\{i\},v\_\{j\}\)\\in\\mathcal\{E\}\_\{\\mathrm\{temp\}\}\\cup\\mathcal\{E\}\_\{\\mathrm\{sem\}\}\\\}\.The reconstructed evidence context is then formed by the anchor turns and their graph neighbors:

Eq=\{uj∣j∈𝒜q∪𝒩𝒢​\(𝒜q\)\}\.\\mathrm\{E\}\_\{q\}=\\\{u\_\{j\}\\mid j\\in\\mathcal\{A\}\_\{q\}\\cup\\mathcal\{N\}\_\{\\mathcal\{G\}\}\(\\mathcal\{A\}\_\{q\}\)\\\}\.This expansion collects turns within a limited graph neighborhood around the anchors, so that the evidence context includes both local dialogue context and semantically related non\-local turns while avoiding excessive irrelevant history\. The selected turns are finally ordered according to their original chronological positions before being passed to the LLM\.

### 4\.4Memory Management

CueMem manages memory through a lightweight design based on cue records, source\-turn pointers, and graph edges\. This design separates offline construction from online reconstruction and supports simple memory operations\.

#### Online and offline workflow\.

CueMem separates offline memory construction from online query\-time reconstruction and answering\. Memory cue extraction and turn graph construction are conducted offline, whereas cue retrieval, context reconstruction, and answer generation are performed online for each incoming query\. This workflow avoids processing the full dialogue history during inference, while still enabling the system to recover rich dialogue evidence through cue\-guided graph expansion\. Moreover, since memory cues contain relatively atomic semantic information, they can be effectively represented by lightweight embedding models, allowing CueMem to balance retrieval efficiency and quality\.

#### Memory maintenance\.

CueMem avoids complex memory maintenance operations by operating on simple cue records, source\-turn pointers, and graph edges, rather than managing consolidated memory units\. ForADD, CueMem extracts cues from newly observed dialogue turns and inserts the corresponding cue records and source\-turn nodes in the background\. Since cues are constructed at the turn level, new memories can be added in near real time, without waiting for a large amount of dialogue history to accumulate for segmentation or compression\. ForDELETE, CueMem can adopt standard forgetting policies, such as time\-based decay or least\-recently\-used removal, to delete obsolete turn nodes and their associated cue records and graph edges\. ForUPDATE, we do not introduce a specialized updating mechanism, instead, updated information is added as new cue records linked to newer source turns\. During generation, the model is instructed to prioritize more recent evidence\. The effectiveness of this simple strategy is further examined in the ablation study\.

## 5Experiments

### 5\.1Experimental Settings

#### Datasets\.

We evaluate CueMem on two long\-term conversational question answering benchmarks: LoCoMo\([Maharana et al\., 2024](https://arxiv.org/html/2609.12354#bib.bib3)\)and LongMemEval\([Wu et al\., 2025](https://arxiv.org/html/2609.12354#bib.bib4)\)\. For LoCoMo, we follow prior work\([Fang et al\., 2026](https://arxiv.org/html/2609.12354#bib.bib5)\)and evaluate on 10 long conversations with 1,540 questions covering single\-hop, multi\-hop, temporal, and open\-domain categories\. For LongMemEval, we use the LongMemEval\-S setting, which contains 500 evaluation questions spanning single\-session user, assistant, and preference questions, as well as multi\-session, temporal\-reasoning, and knowledge\-update questions\. Detailed dataset statistics are provided in the Appendix\.

Table 1:Accuracy results on the LoCoMo dataset\.Table 2:Accuracy results on LongMemEval dataset\.
#### Metrics\.

We report accuracy as the primary evaluation metric\. Following prior work\([Wu et al\., 2025](https://arxiv.org/html/2609.12354#bib.bib4);[Fang et al\., 2026](https://arxiv.org/html/2609.12354#bib.bib5)\), we use an LLM\-as\-judge protocol to compare each generated answer with the reference answer given the corresponding question\. The judge determines whether the prediction correctly answers the question, and accuracy is computed as the proportion of predictions judged correct\.

#### Baselines\.

We compare CueMem with representative long\-term memory systems for conversational agents:

- •Naive RAG \(turn\-level\)retrieves dialogue turns directly from the full conversation history using semantic similarity and uses the retrieved turns as context for answer generation\.
- •MemoryOS\([Kang et al\., 2025](https://arxiv.org/html/2609.12354#bib.bib6)\)organizes conversational memory with an OS\-inspired hierarchical architecture, including short\-term, mid\-term, and long\-term memory units with dynamic updating and retrieval\.
- •Mem0\([Chhikara et al\., 2025](https://arxiv.org/html/2609.12354#bib.bib7)\)builds scalable long\-term memory for production\-ready agents by extracting and maintaining memories from dialogue history through LLM\-guided operations\.
- •A\-Mem\([Xu et al\., 2025](https://arxiv.org/html/2609.12354#bib.bib8)\)constructs structured memory notes and dynamically links related memories following the Zettelkasten principle, enabling agentic memory organization and evolution\.
- •LightMem\([Fang et al\., 2026](https://arxiv.org/html/2609.12354#bib.bib5)\)filters, groups, and consolidates dialogue history through sensory, short\-term, and long\-term memory stages for lightweight memory\-augmented generation\.

#### Implementation details\.

For all methods, we use Llama\-3\.3\-70B\-Instruct as the same LLM backbone for memory extraction, answer generation, and LLM\-as\-judge evaluation\. We adopt all\-MiniLM\-L6\-v2 as a lightweight embedding model for encoding memory and user queries\. Additional results with other LLM and embedding models, as well as more detailed experimental settings, are provided in the Appendix\.

Figure 3:Ablation results on LoCoMo and LongMemEval\.

### 5\.2Main Results

We conduct experiments on LoCoMo and LongMemEval to evaluate the effectiveness of CueMem\. Table[1](https://arxiv.org/html/2609.12354#S5.T1)shows the results on LoCoMo\. CueMem achieves the best overall accuracy, outperforming all baseline methods, and surpassing the strongest baseline by a margin of 5\.26%\. Analyzing performance across question types, CueMem shows clear advantages on single\-hop and temporal questions, improving over the baseline by 4\.40% and 4\.05%, respectively\. This suggests that fine\-grained memory cues are effective for locating specific historical information, while graph\-based expansion helps recover temporally related evidence around the retrieved anchors\. CueMem also performs well on multi\-hop questions, which indicates that the reconstructed context is not limited to the directly matched cues, but can incorporate additional evidence from semantically related or temporally neighboring turns\. For open\-domain questions, CueMem performs slightly worse mainly because these questions require integrating dialogue information with external knowledge, whereas CueMem primarily retrieves and reconstructs evidence from the dialogue history itself, suggesting that explicit external\-knowledge retrieval or higher\-level memory abstraction could further improve its performance on such questions\.

Table[2](https://arxiv.org/html/2609.12354#S5.T2)reports the results on LongMemEval\. CueMem achieves the best overall accuracy, outperforming the strongest baseline by 1\.60%\. This indicates that cue\-guided retrieval and dialogue\-context reconstruction remain effective even in substantially longer conversation histories\. For single\-session questions, CueMem shows strong performance, achieving the best results on both single\-user and single\-preference categories and remaining competitive on single\-assistant questions\. This suggests that fine\-grained memory cues are effective for capturing user\-specific facts and preferences within localized interaction contexts\. Since single\-session questions are common in practical conversational agent scenarios, accurately answering such questions is important for improving the user experience and maintaining coherent personalized interactions\. CueMem also achieves the best performance on knowledge\-update questions, even though it does not introduce a specializedUPDATEoperation\. This may be because CueMem leverages fine\-grained cues to accurately and comprehensively reconstruct query\-relevant evidence from the original dialogue while preserving temporal order\. Even when the reconstructed context contains conflicting historical information, modern LLMs can often identify and ignore outdated evidence when properly instructed\. This result suggests that explicit memory rewriting may not always be necessary for handling knowledge updates\. When relevant dialogue evidence is accurately reconstructed and temporally ordered, a sufficiently capable LLM can often resolve conflicts between old and updated information during answer generation\.

Across both datasets, CueMem demonstrates consistently strong performance across different long dialogue QA settings\. In contrast, most baselines tend to perform well on only one dataset\. For example, LightMem achieves stronger results on LongMemEval than on LoCoMo, while several other baselines obtain reasonable performance on LoCoMo but degrade substantially on LongMemEval\. This contrast suggests that CueMem is less sensitive to dataset\-specific characteristics and can generalize more consistently across conversations with different lengths, structures, and question distributions\.

### 5\.3Ablation Study

In this section, we conduct two ablation studies on LoCoMo and LongMemEval to evaluate the effectiveness of turn graph expansion and the priority instruction, respectively\.

On LoCoMo, we disable graph expansion while keeping the same cue extraction and retrieval components\. The ablated variant therefore answers questions using only the anchor turns associated with the retrieved cues as context\. As shown in Figure[3](https://arxiv.org/html/2609.12354#S5.F3)\(a\), removing the turn graph reduces the overall accuracy from 81\.1% to 71\.4%, with the largest drops on single\-hop and multi\-hop questions\. This suggests that the retrieved anchor turns alone often do not contain sufficient evidence, and that graph expansion is important for recovering temporally adjacent and semantically related supporting turns\.

As discussed above, CueMem does not perform an explicitUPDATEoperation, yet it still achieves strong performance on knowledge\-update questions\. We attribute this result largely to the priority instruction used during answer generation\. In practice, this is implemented by adding a simple prompt instruction: “Prioritize the information most recent and closest to the question time”\. To verify whether this simple instruction is indeed effective, we remove it while keeping the retrieved evidence context unchanged\. As shown in Figure[3](https://arxiv.org/html/2609.12354#S5.F3)\(b\), removing the priority instruction lowers the overall accuracy from 75\.2% to 72\.4%, and the accuracy on knowledge\-update questions drops from 87\.2% to 80\.8%\. In contrast, several non\-update categories remain unchanged or show only minor differences\. Temporal questions also show a noticeable accuracy drop of 6\.8%, indicating that the priority instruction not only helps the model identify updated facts, but also guides it toward more accurate temporal reasoning\. These ablation results suggest that memory systems can handle knowledge updates without relying on complex update operations\. Instead, when relevant evidence is accurately reconstructed, the inherent reasoning capability of the LLM, combined with a simple priority instruction, can effectively resolve update\-sensitive questions\.

Table 3:Efficiency and effectiveness comparison between the full\-history LLM setting and CueMem\. All numbers are averaged per query\. Tokens denote the number of input tokens provided to the answer\-generation LLM\. Recon\. denotes the query\-time context reconstruction cost, including query encoding, cue retrieval, and graph expansion\. Infer\. denotes the LLM answer\-generation latency\.
### 5\.4Efficiency Analysis

As LLMs continue to support increasingly long context windows, directly feeding the entire dialogue history into the LLM has become a simple and seemingly attractive solution for long\-term conversational question answering\. In this section, we compare CueMem with the full\-history LLM setting in terms of both effectiveness and query\-time efficiency, focusing on answer accuracy, token consumption, and answering latency, with CueMem using a reconstructed context of approximately 2K tokens\.

Table[3](https://arxiv.org/html/2609.12354#S5.T3)shows that CueMem is substantially more efficient than the full\-history LLM setting while achieving better accuracy\. On LoCoMo, CueMem consumes only about 10% of the input tokens used by the full\-history setting while achieving higher accuracy\. Meanwhile, it reduces the overall query latency by 42\.5%\. On LongMemEval, this advantage becomes even more pronounced\. Compared with the full\-history setting, CueMem uses less than 2% of the input tokens, achieves a 27\.2% absolute improvement in accuracy, and reduces the overall query latency by 87\.4%\.

These results suggest that long\-context LLMs may work reasonably well with moderately long histories, but can degrade when the input approaches the maximum context window \(such as 128K limit for Llama\-3\.3\)\. This degradation may be caused by excessive irrelevant context, noise, or the lost\-in\-the\-middle effect\. Moreover, the full\-history strategy inevitably increases inference cost and latency as the conversation grows\. In contrast, CueMem mitigates these issues with only a small context reconstruction overhead, substantially reducing token consumption and latency while improving answer accuracy\. This low overhead is enabled by CueMem’s fine\-grained cues, whose atomic semantics can be effectively encoded by a lightweight 22M\-parameter embedding model \(all\-MiniLM\-L6\-v2\) with 384\-dimensional representations, making the retrieval module lightweight enough for resource\-constrained settings\.

## 6Conclusion

We presented CueMem, a lightweight and effective long\-term memory framework for conversational agents\. CueMem follows a simple cue\-to\-anchor\-to\-context pipeline: it treats extracted memories as fine\-grained retrieval cues rather than self\-contained evidence, links them to source turns, and reconstructs query\-relevant dialogue context through expansion over a turn graph\. This design keeps memory management simple while retaining direct links to the original dialogue evidence\. Experiments on LoCoMo and LongMemEval show that CueMem consistently outperforms representative long\-term memory baselines while reducing query\-time token consumption and latency compared with the full\-history setting\. These results suggest that cue\-guided context reconstruction offers a simple and scalable strategy for long\-term memory in conversational agents\.

## References

- Achiamet al\.\(2023\)J\. Achiam, S\. Adler, S\. Agarwal, L\. Ahmad, I\. Akkaya, F\. L\. Aleman, D\. Almeida, J\. Altenschmidt, S\. Altman, S\. Anadkat,et al\.Gpt\-4 technical report\.arXiv preprint arXiv:2303\.08774\.Cited by:[§1](https://arxiv.org/html/2609.12354#S1.p1.1)\.
- Chhikaraet al\.\(2025\)P\. Chhikara, D\. Khant, S\. Aryan, T\. Singh, and D\. YadavMem0: building production\-ready ai agents with scalable long\-term memory\.arXiv preprint arXiv:2504\.19413\.Cited by:[§2](https://arxiv.org/html/2609.12354#S2.p2.1),[3rd item](https://arxiv.org/html/2609.12354#S5.I1.i3.p1.1),[Table 1](https://arxiv.org/html/2609.12354#S5.T1.2.5.1)\.
- Conway and Rubin \(1993\)M\. Conway and D\. RubinThe structure of autobiographical memory\.InTheories of Memory,pp\. 103–138\.Cited by:[§1](https://arxiv.org/html/2609.12354#S1.p2.1),[§2](https://arxiv.org/html/2609.12354#S2.p3.1)\.
- Conway and Pleydell\-Pearce \(2000\)M\. A\. Conway and C\. W\. Pleydell\-PearceThe construction of autobiographical memories in the self\-memory system\.Psychological Review107\(2\),pp\. 261–288\.Cited by:[§1](https://arxiv.org/html/2609.12354#S1.p2.1),[§2](https://arxiv.org/html/2609.12354#S2.p3.1)\.
- Edgeet al\.\(2024\)D\. Edge, H\. Trinh, N\. Cheng, J\. Bradley, A\. Chao, A\. Mody, S\. Truitt, D\. Metropolitansky, R\. O\. Ness, and J\. LarsonFrom local to global: a graph rag approach to query\-focused summarization\.arXiv preprint arXiv:2404\.16130\.Cited by:[§2](https://arxiv.org/html/2609.12354#S2.p1.1)\.
- Fanget al\.\(2026\)J\. Fang, X\. Deng, H\. Xu, Z\. Jiang, Y\. Tang, Z\. Xu, S\. Deng, Y\. Yao, M\. Wang, S\. Qiao, H\. Chen, and N\. ZhangLightMem: lightweight and efficient memory\-augmented generation\.InThe Fourteenth International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2609.12354#S1.p2.1),[§2](https://arxiv.org/html/2609.12354#S2.p2.1),[5th item](https://arxiv.org/html/2609.12354#S5.I1.i5.p1.1),[§5\.1](https://arxiv.org/html/2609.12354#S5.SS1.SSS0.Px1.p1.1),[§5\.1](https://arxiv.org/html/2609.12354#S5.SS1.SSS0.Px2.p1.1),[Table 1](https://arxiv.org/html/2609.12354#S5.T1.2.7.1)\.
- Gaoet al\.\(2023\)Y\. Gao, Y\. Xiong, X\. Gao, K\. Jia, J\. Pan, Y\. Bi, Y\. Dai, J\. Sun, H\. Wang,et al\.Retrieval\-augmented generation for large language models: a survey\.arXiv preprint arXiv:2312\.109972\(1\),pp\. 32\.Cited by:[§2](https://arxiv.org/html/2609.12354#S2.p1.1)\.
- Guoet al\.\(2024\)Z\. Guo, L\. Xia, Y\. Yu, T\. Ao, and C\. HuangLightrag: simple and fast retrieval\-augmented generation\.arXiv preprint arXiv:2410\.057792\(3\)\.Cited by:[§2](https://arxiv.org/html/2609.12354#S2.p1.1)\.
- Gutiérrezet al\.\(2024\)B\. J\. Gutiérrez, Y\. Shu, Y\. Gu, M\. Yasunaga, and Y\. SuHipporag: neurobiologically inspired long\-term memory for large language models\.Advances in neural information processing systems37,pp\. 59532–59569\.Cited by:[§2](https://arxiv.org/html/2609.12354#S2.p1.1)\.
- Gutiérrezet al\.\(2025\)B\. J\. Gutiérrez, Y\. Shu, W\. Qi, S\. Zhou, and Y\. SuFrom rag to memory: non\-parametric continual learning for large language models\.InInternational Conference on Machine Learning,pp\. 21497–21515\.Cited by:[§2](https://arxiv.org/html/2609.12354#S2.p1.1)\.
- Holland and Kensinger \(2010\)A\. C\. Holland and E\. A\. KensingerEmotion and autobiographical memory\.Physics of Life Reviews7\(1\),pp\. 88–131\.Cited by:[§1](https://arxiv.org/html/2609.12354#S1.p2.1),[§2](https://arxiv.org/html/2609.12354#S2.p3.1)\.
- Kanget al\.\(2025\)J\. Kang, M\. Ji, Z\. Zhao, and T\. BaiMemory os of ai agent\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 25961–25970\.Cited by:[§2](https://arxiv.org/html/2609.12354#S2.p2.1),[2nd item](https://arxiv.org/html/2609.12354#S5.I1.i2.p1.1),[Table 1](https://arxiv.org/html/2609.12354#S5.T1.2.4.1)\.
- Lewiset al\.\(2020\)P\. Lewis, E\. Perez, A\. Piktus, F\. Petroni, V\. Karpukhin, N\. Goyal, H\. Küttler, M\. Lewis, W\. Yih, T\. Rocktäschel,et al\.Retrieval\-augmented generation for knowledge\-intensive nlp tasks\.Advances in neural information processing systems33,pp\. 9459–9474\.Cited by:[§2](https://arxiv.org/html/2609.12354#S2.p1.1)\.
- Liuet al\.\(2024\)N\. F\. Liu, K\. Lin, J\. Hewitt, A\. Paranjape, M\. Bevilacqua, F\. Petroni, and P\. LiangLost in the middle: how language models use long contexts\.Transactions of the association for computational linguistics12,pp\. 157–173\.Cited by:[§1](https://arxiv.org/html/2609.12354#S1.p1.1),[§2](https://arxiv.org/html/2609.12354#S2.p1.1)\.
- Maharanaet al\.\(2024\)A\. Maharana, D\. Lee, S\. Tulyakov, M\. Bansal, F\. Barbieri, and Y\. FangEvaluating very long\-term conversational memory of llm agents\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 13851–13870\.Cited by:[§1](https://arxiv.org/html/2609.12354#S1.p4.1),[§5\.1](https://arxiv.org/html/2609.12354#S5.SS1.SSS0.Px1.p1.1)\.
- Packeret al\.\(2023\)C\. Packer, S\. Wooders, K\. Lin, V\. Fang, S\. G\. Patil, I\. Stoica, and J\. E\. GonzalezMemGPT: towards llms as operating systems\.arXiv preprint arXiv:2310\.08560\.Cited by:[§2](https://arxiv.org/html/2609.12354#S2.p2.1)\.
- Panet al\.\(2025\)Z\. Pan, Q\. Wu, H\. Jiang, X\. Luo, H\. Cheng, D\. Li, Y\. Yang, C\. Lin, H\. V\. Zhao, L\. Qiu, and J\. GaoSeCom: on memory construction and retrieval for personalized conversational agents\.InThe Thirteenth International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2609.12354#S1.p2.1)\.
- Parket al\.\(2023\)J\. S\. Park, J\. O’Brien, C\. J\. Cai, M\. R\. Morris, P\. Liang, and M\. S\. BernsteinGenerative agents: interactive simulacra of human behavior\.InProceedings of the 36th annual acm symposium on user interface software and technology,pp\. 1–22\.Cited by:[§2](https://arxiv.org/html/2609.12354#S2.p2.1)\.
- Wuet al\.\(2025\)D\. Wu, H\. Wang, W\. Yu, Y\. Zhang, K\. Chang, and D\. YuLongMemEval: benchmarking chat assistants on long\-term interactive memory\.InThe Thirteenth International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2609.12354#S1.p4.1),[§5\.1](https://arxiv.org/html/2609.12354#S5.SS1.SSS0.Px1.p1.1),[§5\.1](https://arxiv.org/html/2609.12354#S5.SS1.SSS0.Px2.p1.1)\.
- Xuet al\.\(2025\)W\. Xu, Z\. Liang, K\. Mei, H\. Gao, J\. Tan, and Y\. ZhangA\-mem: agentic memory for llm agents\.Advances in Neural Information Processing Systems38,pp\. 17577–17604\.Cited by:[§2](https://arxiv.org/html/2609.12354#S2.p2.1),[4th item](https://arxiv.org/html/2609.12354#S5.I1.i4.p1.1),[Table 1](https://arxiv.org/html/2609.12354#S5.T1.2.6.1)\.
- Yanget al\.\(2025\)A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv,et al\.Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.Cited by:[§1](https://arxiv.org/html/2609.12354#S1.p1.1)\.
- Zhonget al\.\(2024\)W\. Zhong, L\. Guo, Q\. Gao, H\. Ye, and Y\. WangMemorybank: enhancing large language models with long\-term memory\.InProceedings of the AAAI conference on artificial intelligence,Vol\.38,pp\. 19724–19731\.Cited by:[§1](https://arxiv.org/html/2609.12354#S1.p2.1),[§2](https://arxiv.org/html/2609.12354#S2.p2.1)\.

Similar Articles

MemTrain: Self-Supervised Context Memory Training

arXiv cs.CL

MemTrain proposes a self-supervised training framework that uses masked reconstruction and intermediate memory recall proxy tasks on Wikipedia corpora to enhance LLM agents' context memory, achieving up to 17.67 point gains on downstream memory-intensive QA benchmarks.