EnSIMem: 用于长期代理记忆的实体结构索引
摘要
EnSIMem是一种用于长期AI代理的实体结构化记忆架构,通过实体-属性索引将交互组织成事件,在基准测试中实现高召回准确率。
查看缓存全文
缓存时间: 2026/09/24 09:19
# EnSIMem: Entity-Structured Indexing for Long-Term Agent Memory
Source: [https://arxiv.org/html/2609.27279](https://arxiv.org/html/2609.27279)
Xing FanAffiliation:AmazonEmail:[xfan31@illinois\.edu](mailto:)Xinyi FanAffiliation:University of Illinois Urbana\-ChampaignEmail:[xie39@illinois\.edu](mailto:)Chenlei GuoAffiliation:AmazonEmail:[hanj@illinois\.edu](mailto:)Yixuan XieAffiliation:University of Illinois Urbana\-ChampaignEmail:[fanxing@amazon\.com](mailto:)Jiawei HanAffiliation:University of Illinois Urbana\-ChampaignEmail:[guochenl@amazon\.com](mailto:)
###### Abstract
An agent that interacts with users over long periods must recall facts, preferences, events, and changes from a continuously growing interaction history\. Existing memory systems often compress interactions into generic summaries or retrieve anonymous text chunks, making it difficult for an agent to identify the correct entity, property, and supporting evidence\. We present EnSIMem, an entity\-structured long\-term memory architecture for an agent\. During offline construction, the system organizes interactions into theme\-coherent episodes and builds dialogue\-grounded index entries of the form\[entity\]\[entity\_type\]\[property:value\]\[\\textit\{entity\}\]\[\\textit\{entity\\\_type\}\]\[\\textit\{property\}:\\textit\{value\}\]\. Each entry preserves its source turns, temporal information, and available multimodal fields\. During online interaction, the agent’s request is decomposed into evidence requirements whose properties are aligned with the memory index\. Entity\-property lookup and adaptive retrieval then collect the evidence needed for point, temporal, compositional, and aggregation reasoning\. The agent generates its response from the preserved source evidence rather than from lossy memory summaries\. On long\-term agent\-memory benchmarks, EnSIMem achieves high answer accuracy while maintaining compact contexts and favorable online efficiency\. These results show that entity\-structured indexing and episode\-level provenance provide a reliable foundation for long\-term memory in agents\. The code of our model is available at[https://github\.com/RamonMeng/EnSIMem](https://github.com/RamonMeng/EnSIMem)\.
## 1Introduction
An agent deployed over along termmust accumulate and reuse information from its past interactions\. This information may include user preferences, personal facts, plans, commitments, events, feedback, and decisions that influence future actions\. For an agent, memory is therefore not simply a larger context window: it is a persistent interface between past experience and current reasoning, as reflected by memory\-stream, tiered\-memory, and evolving\-memory architectures\([Park et al\., 2023](https://arxiv.org/html/2609.27279#bib.bib22);[Packer et al\., 2023](https://arxiv.org/html/2609.27279#bib.bib23);[Xu et al\., 2025](https://arxiv.org/html/2609.27279#bib.bib29);[Kang et al\., 2025](https://arxiv.org/html/2609.27279#bib.bib30)\)\. The memory system must help the agent recover the right evidence at the right time while preserving enough context for the agent to interpret that evidence correctly\.
Long\-term agent memory is particularly challenging because relevant information is distributed across many interactions and may be expressed in different ways\. This challenge is central to long\-term conversational\-memory benchmarks and interactive memory evaluations\([Maharana et al\., 2024](https://arxiv.org/html/2609.27279#bib.bib18);[Wu et al\., 2024](https://arxiv.org/html/2609.27279#bib.bib21)\)\. A user may describe an event with a concrete action while a later request refers to the same event using a more general expression\. Important information may also be introduced through pronouns, temporal references, correlations, or multimodal observations\. For example, an interaction may state that a person traveled to Chicago by helicopter, while a later request asks how that person traveled to Chicago\. A useful memory system must connect these expressions without collapsing the underlying event into an uninformative category such as “activity” or “travel\.” It must also retain the original evidence so that the agent can distinguish an explicit fact from an inference\.
Existing approaches expose a tension between scalability and specificity\. Providing the complete interaction history gives the agent access to all evidence, but causes context growth, higher latency, and distraction effects\([Liu et al\., 2024](https://arxiv.org/html/2609.27279#bib.bib1);[Du et al\., 2025](https://arxiv.org/html/2609.27279#bib.bib16);[Wang et al\., 2024](https://arxiv.org/html/2609.27279#bib.bib2);[Bai et al\., 2025](https://arxiv.org/html/2609.27279#bib.bib3)\)\. Summarization reduces the context size but can remove details, temporal qualifiers, and provenance\. Conventional retrieval\-augmented generation systems retrieve anonymous text chunks or vector\-nearest memories\([Lewis et al\., 2020](https://arxiv.org/html/2609.27279#bib.bib4);[Karpukhin et al\., 2020](https://arxiv.org/html/2609.27279#bib.bib7);[Izacard et al\., 2022](https://arxiv.org/html/2609.27279#bib.bib8);[Khattab and Zaharia, 2020](https://arxiv.org/html/2609.27279#bib.bib9)\); more recent systems use iterative or structured retrieval to improve multi\-step reasoning\([Asai et al\., 2024](https://arxiv.org/html/2609.27279#bib.bib5);[Trivedi et al\., 2023](https://arxiv.org/html/2609.27279#bib.bib6);[Edge et al\., 2024](https://arxiv.org/html/2609.27279#bib.bib10)\)\. These representations are compact, but they do not explicitly indicate which entity, property, or event makes a memory relevant\. EnSIMem addresses this tension by using structure to identify the right evidence and the original dialogue to interpret it: a large interaction history is converted into a small, query\-specific, evidence\-rich buffer instead of being passed wholesale to the agent\.
We present EnSIMem, an entity\-structure indexing\-based long\-term memory architecture for agentic systems, building on entity\-structured retrieval and structure\-augmented reasoning ideas\([Meng et al\., 2026](https://arxiv.org/html/2609.27279#bib.bib20);[Parekh et al\., 2025](https://arxiv.org/html/2609.27279#bib.bib14)\)\. Figure[1](https://arxiv.org/html/2609.27279#S1.F1)summarizes the overall EnSIMem architecture\. The central design principle is to use structure as an address system for memory, rather than as a lossy replacement for the original interaction\. During offline construction, EnSIMem organizes conversations into theme\-coherent episodes and builds dialogue\-grounded index records of the form\[entity\]\[entity\_type\]\[property:value\]\[\\textit\{entity\}\]\[\\textit\{entity\\\_type\}\]\[\\textit\{property\}:\\textit\{value\}\]\. The entity may be a person, object, or salient event\. The property is selected at an intermediate level of granularity: broad enough to connect paraphrases, but specific enough to preserve the identity of the underlying relation or event\. When available, a finer\-grained value records the concrete realization of that property\. Every record remains linked to its source turns, timestamps, and multimodal fields\.
Upon interaction, the query is decomposed into explicit evidence requirements\. This follows the broader use of iterative query decomposition and agentic query rewriting for complex retrieval tasks\([Trivedi et al\., 2023](https://arxiv.org/html/2609.27279#bib.bib6);[Shankar et al\., 2024](https://arxiv.org/html/2609.27279#bib.bib15)\)\. The query planner extracts entities and properties at the same granularity as the memory index and identifies whether the query requires point evidence, temporal comparison, compositional reasoning, or aggregation\. Structured entity\-property lookup is combined with dense fallback retrieval when the wording differs from the indexed property\. Retrieval then proceeds adaptively until the requirements are sufficiently covered\. This allows the system to retrieve a small candidate set for a point query while preserving exhaustive coverage for a query that requires counting or comparing multiple events\.
The final response is generated from the retrieved source evidence rather than from index entries alone, consistent with evidence\-grounded graph and structure\-aware generation approaches\([He et al\., 2024](https://arxiv.org/html/2609.27279#bib.bib12);[Hu et al\., 2024](https://arxiv.org/html/2609.27279#bib.bib13)\)\. This separation gives the agent both efficient access and evidentiary grounding: the index identifies where to look, while the original episode provides the context needed for interpretation and answer synthesis\. It also makes the memory pipeline inspectable\. Each answer can be traced from a query requirement, through an entity\-property access path, to the dialogue turns and multimodal observations that support it\.
EnSIMem separates query\-independent memory construction from query\-dependent recall\. The offline representation can therefore be reused across many future queries, while the online stage spends computation only on the evidence required by the current query\. On long\-term conversational memory benchmarks, EnSIMem achieves high answer accuracy while maintaining compact retrieval contexts and favorable online efficiency\. The results suggest that theme\-coherent episodes, dialogue\-grounded surface\-property indexing, and requirement\-aware retrieval provide a practical foundation for reliable long\-term memory in agents\.
Our contributions are:
1. 1\.Theme\-coherent episodic memory construction\.We organize long interactions into theme\-coherent, contiguous episodes that preserve local context, temporal order, and provenance\.
2. 2\.Dialogue\-grounded entity\-property indexing\.We index originally mentioned entities and salient events with structured records linked to their original dialogue turns and multimodal observations\.
3. 3\.Granularity\-controlled property alignment\.We align corpus and query properties at an intermediate level that supports paraphrase matching without erasing event\-specific distinctions\.
4. 4\.Requirement\-aware evidence localization\.We decompose agent requests into explicit evidence requirements and retrieve memories through structured entity\-property matching\.
5. 5\.Query\-type\-aware adaptive retrieval\.We distinguish point queries from temporal, compositional, and aggregation queries and allocate evidence budgets appropriate to their reasoning requirements\.
6. 6\.An evidence\-rich short\-context buffer for agent response generation\.We convert long interaction histories into compact, query\-specific reasoning buffers that retain relevant source evidence while filtering unrelated noise, without additional training or supervision\. This methodology leads to state\-of\-the\-art results, as shown in our experiments\.

Figure 1:Overview of EnSIMem\. Offline memory construction preserves multimodal dialogue, partitions it into theme\-coherent episodes, extracts dialogue\-grounded entity\-property records, and stores an entity\-property index linked to the original episodes\. Online, the agent decomposes the request, performs requirement\-aware retrieval with structured matching, expands the selected episode into source evidence, and generates a grounded response\.
## 2Related Work
#### Memory architectures for long\-running agents\.
Prior agents use memory streams and reflection\([Park et al\., 2023](https://arxiv.org/html/2609.27279#bib.bib22)\), tiered context management\([Packer et al\., 2023](https://arxiv.org/html/2609.27279#bib.bib23)\), continual personalization\([Zhong et al\., 2023](https://arxiv.org/html/2609.27279#bib.bib24)\), verbal experience\([Shinn et al\., 2023](https://arxiv.org/html/2609.27279#bib.bib25)\), or reusable skills\([Wang et al\., 2023](https://arxiv.org/html/2609.27279#bib.bib26)\)\. Recent systems add temporal graphs, consolidated memories, evolving notes, and hierarchical stores\([Rasmussen et al\., 2025](https://arxiv.org/html/2609.27279#bib.bib27);[Chhikara et al\., 2025](https://arxiv.org/html/2609.27279#bib.bib28);[Xu et al\., 2025](https://arxiv.org/html/2609.27279#bib.bib29);[Kang et al\., 2025](https://arxiv.org/html/2609.27279#bib.bib30)\)\. EnSIMem complements these approaches with source\-grounded entity\-property access for factual and relational recall\.
#### Entity\-structured retrieval for long documents\.
Dense retrieval and graph\-aware methods support general\-purpose and multi\-hop access\([Lewis et al\., 2020](https://arxiv.org/html/2609.27279#bib.bib4);[Karpukhin et al\., 2020](https://arxiv.org/html/2609.27279#bib.bib7);[Khattab and Zaharia, 2020](https://arxiv.org/html/2609.27279#bib.bib9);[Edge et al\., 2024](https://arxiv.org/html/2609.27279#bib.bib10);[Jiménez Gutiérrez et al\., 2024](https://arxiv.org/html/2609.27279#bib.bib11)\)\. EnSI\-RAG introduces entity\-structure indexing for long\-document question answering\([Meng et al\., 2026](https://arxiv.org/html/2609.27279#bib.bib20)\); EnSIMem extends this idea to persistent agent memory by linking dialogue\-grounded entity\-property records to theme\-coherent episodes and query\-aware evidence budgets\.
#### Structured agent memory\.
Memora separates rich values from lightweight abstractions and cue anchors, while EnSIMem retains the original source turns and uses explicit entity\-property alignment to localize evidence\([Xia et al\., 2026](https://arxiv.org/html/2609.27279#bib.bib19)\)\. Detailed benchmark definitions and baseline provenance are deferred to Appendix[A](https://arxiv.org/html/2609.27279#A1)\.
## 3EnSIMem
### 3\.1System Objective and Design Principle
EnSIMem is an index\-structured long\-term memory layer for an agent that can interact with systems/users over an extended period of time\. Given an interaction corpus𝒟\\mathcal\{D\}and a new agent requestqq, the system must identify the relevant past evidence and provide it to the agent for reasoning and response generation\.
We separate memory construction from memory access:
ℳ=Fmemory\(𝒟\),y^=G\(q,R\(q,ℳ\)\),\\mathcal\{M\}=F\_\{\\mathrm\{memory\}\}\(\\mathcal\{D\}\),\\qquad\\hat\{y\}=G\\\!\\left\(q,R\\\!\\left\(q,\\mathcal\{M\}\\right\)\\right\),\(1\)whereℳ\\mathcal\{M\}is a persistent, query\-independent memory representation of corpus𝒟\\mathcal\{D\}via its memory construction processFmemoryF\_\{\\mathrm\{memory\}\},RRis the query\-dependent retrieval process upon queryqq, andGGis the agent’s response\-generation process, which generatesy^\\hat\{y\}\.
The central design principle is to use an index structure as an address system for evidence, rather than a lossy replacement for the original interaction\. Indexed structured records make it possible to identify which entity, event, and property are relevant, while the original dialogue remains available for interpretation and answer synthesis\.
### 3\.2Theme\-Coherent Episodic Memory
The interaction corpus is an ordered collection of sessions containing text, speakers, timestamps, and optional multimodal observations\. EnSIMem groups neighboring turns into contiguous, theme\-coherent episodes rather than treating every turn or an entire session as a retrieval unit\. Because the partition is query\-independent and unsupervised, we use a fixed semantic policy: turns remain together when they share an entity, event, goal, or unresolved conversational reference, and a boundary is introduced only when a new topic or task persists\. A speaker change alone does not trigger a split, and multimodal fields remain attached to the turn that introduced them\.
At a high level, the query\-independent partition can be written as
Π\(Sn\)=\(en,1,…,en,mn\),⋃j=1mnen,j=Sn,en,j∩en,k=∅\(j≠k\),\\Pi\(S\_\{n\}\)=\(e\_\{n,1\},\\ldots,e\_\{n,m\_\{n\}\}\),\\qquad\\bigcup\_\{j=1\}^\{m\_\{n\}\}e\_\{n,j\}=S\_\{n\},\\quad e\_\{n,j\}\\cap e\_\{n,k\}=\\varnothing\\;\(j\\neq k\),\(2\)where eachen,je\_\{n,j\}is a contiguous, theme\-coherent portion of sessionSnS\_\{n\}\. The equation emphasizes coverage and non\-overlap, while the semantic policy determines the useful granularity\. A theme is a short description of what the turns are jointly about, rather than a fixed ontology or a query\-specific label\. The policy keeps follow\-up turns and references with the material they depend on, but separates a sustained change of topic, event, or goal\.
This construction is a middle ground between individual turns and complete sessions: per\-turn units can lose anaphoric context, whereas whole\-session units can mix unrelated topics and add noise\. Because segmentation is unsupervised, we do not claim a unique correct boundary for every conversation; nearby choices at the same semantic scale should preserve the same underlying evidence and have similar behavior\. The episode\-granularity ablation in Section 4 tests this trade\-off by comparing theme\-coherent, per\-turn, and per\-session units\.
### 3\.3Dialogue\-Grounded Entity–Property Index
For each episodeee, EnSIMem extracts explicitly mentioned entities and salient events, including people, objects, places, activities, and named events\. It then associates each entity with properties or relations supported by the local dialogue\.
The extractor produces records of the form
r=⟨x,t,p,v,e,π⟩,r=\\langle x,t,p,v,e,\\pi\\rangle,\(3\)wherexxis an entity,ttits type,ppa property or relation,vvan optional concrete value,eethe containing episode, andπ\\pia provenance link to the supporting evidence\. The record is an address for finding the episode, not a replacement for it\.
The retrieval\-facing index is therefore written as:
\[x\]\[t\]\[p:v\]⟶\{𝑒𝑝𝑖𝑠𝑜𝑑𝑒\_𝑖𝑑\}\.\[x\]\[t\]\[p:v\]\\longrightarrow\\\{\\mathit\{episode\\\_id\}\\\}\.\(4\)The value is retained when the episode states a concrete realization of the property and is empty when no value is available\. For example, a discussion of Caroline researching adoption agencies yields the index handle
\[Caroline\]\[person\]\[research:adoption agencies\]⟶\{e\}\.\[\\text\{Caroline\}\]\[\\text\{person\}\]\[\\text\{research\}:\\text\{adoption agencies\}\]\\longrightarrow\\\{e\\\}\.\(5\)At query time, this handle identifies a relevant episode; the agent then reads the preserved source context to recover the answer and its qualifications\. The index remains compact and reusable while the source retains temporal, conversational, and multimodal information\.
### 3\.4Granularity\-Controlled Property Alignment
A central requirement of EnSIMem is that properties extracted from the corpus and properties extracted from an agent request must havecompatible granularity\.
A property should be broad enough to support paraphrase or similarity match, but specific enough to preserve the identity of the underlying event or relation\. An overly narrow property reduces recall because semantically equivalent expressions fail to match\. But an overly broad property fails to support discriminative selection\. For example, mappingresearch,attending a class, andplaying footballto a generic propertyactivitymakes it impossible to identify which activity it is about\.
EnSIMem treats property extraction as a controlled abstraction process\. The property captures the semantic relation needed for retrieval, while the value preserves the concrete realization when it is present\. The index does not rely on a fixed, dataset\-specific ontology\. Instead, corpus properties and request properties are generated under the principle ofcompatible semantic granularityand compared through compatibility rather than exact string match\.
This alignment allows expressions such as “took a helicopter to Chicago” and “traveled to Chicago” to share a travel\-related retrieval property while preserving the specific transportation detail in the value\. At the same time, unrelated actions remain distinguishable because they are not collapsed into a universal property \(e\.g\.,activity\)\.
### 3\.5Requirement\-Aware Query Planning
For a requestqq, the query planner extracts a complete set of atomic retrieval requirements from the question\. We denote this requirement set byH\(q\)H\(q\):
H\(q\)=\{h1,…,hK\},hi=\(xi,ti,pi,vi,ci,ρi\),H\(q\)=\\\{h\_\{1\},\\ldots,h\_\{K\}\\\},\\qquad h\_\{i\}=\(x\_\{i\},t\_\{i\},p\_\{i\},v\_\{i\},c\_\{i\},\\rho\_\{i\}\),\(6\)where each requirement identifies a target entityxix\_\{i\}, its typetit\_\{i\}, an intermediate\-granularity propertypip\_\{i\}, an optional valueviv\_\{i\}or conditioncic\_\{i\}, and the roleρi\\rho\_\{i\}that the evidence will play in answering the request\. The planner also classifies the reasoning type of the request\. EnSIMem distinguishes, for example:
- •point questions, which require one or a small number of relevant episodes;
- •temporal questions, which require evidence ordered or compared by time;
- •compositional questions, which require multiple related properties or entities; and
- •aggregation and frequency questions, which require collecting all relevant episodes before counting or comparing them\.
This classification determines the evidence budget before retrieval begins\. A frequency question such as “How often did Sam get health checkups?” cannot be treated as an ordinary top\-kkquestion, because the answer may require counting multiple separate episodes even when no single episode states the final frequency\.
### 3\.6Requirement\-Coverage Retrieval
For each requirement, structured matching first searches the entity\-property index for records compatible in entity, type, property, value, and conditions\. Matching records identify candidate episodes\. Letℰ\(q\)\\mathcal\{E\}\(q\)denote the deduplicated set of candidate episodes returned for queryqq\. IfIIdenotes the structured index andTTits textual representation, then
ℰ\(q\)=Dedup\(Structured\(I,H\(q\)\)∪Dense\(T,H\(q\)\)\)\.\\mathcal\{E\}\(q\)=\\operatorname\{Dedup\}\\\!\\left\(\\operatorname\{Structured\}\(I,H\(q\)\)\\;\\cup\\;\\operatorname\{Dense\}\(T,H\(q\)\)\\right\)\.\(7\)
The structured path is deliberately used as an address system: it can identify an episode even when the useful answer is stated across several nearby turns\. When the structured path is incomplete or the wording differs from the indexed property, dense retrieval over the textual index representation supplies additional candidates\. The two paths are complementary: the index provides a precise address, while dense retrieval protects recall\.
The resulting candidates are merged and deduplicated by episode identity, and the system tracks which requirements are supported by the evidence collected so far\. Point questions can stop after a small set covers the answer\. Temporal and compositional questions continue until their comparisons or reasoning bridges are supported\. Aggregation and frequency questions use an exhaustive stopping condition: all necessary episodes are retrieved before generation\. Stopping is therefore determined by the query’s evidence requirements rather than by an arbitrary numeric similarity threshold\.
### 3\.7Evidence\-Grounded Agent Response
After retrieval, selected episodes are expanded back into source evidence, including turn text, speaker, time, and available multimodal fields\. IfΓ\(ℰ\(q\)\)\\Gamma\(\\mathcal\{E\}\(q\)\)denotes this expansion, the agent response is conceptually
y^=G\(q,H\(q\),Γ\(ℰ\(q\)\)\)\.\\hat\{y\}=G\\\!\\left\(q,H\(q\),\\Gamma\(\\mathcal\{E\}\(q\)\)\\right\)\.\(8\)HereGGis the response\-generation function: it receives the request, its retrieval requirements, and the expanded source evidence, and produces the grounded answery^\\hat\{y\}\. The agent receives this compact evidence buffer together with the request and its requirements, and generates the answer from the source rather than from index records alone\. Explicit facts are answered directly, uncertainty is acknowledged when evidence is incomplete, temporal wording preserves the precision of the source, and aggregation is computed across the retrieved episodes\.
For example, for the request “What did Caroline research?” the planner identifies the entity Caroline and the propertyresearch\. The index points to the episode in which she says that she is researching adoption agencies, and the agent expands that episode and answers “adoption agencies\.” The same process preserves an image caption, image\-associated query, or image URL when it contains information not present in the dialogue text\.
### 3\.8Separation of Offline memory construction and online agent request processing
Our design separates offline persistent construction from online request\-dependent access\. Offline processing turns the agent interaction corpus into reusable, provenance\-linked indexed memory\. For each request, the online stage selects a requirement\-covered set of episodes and expands it into a compact evidence buffer for agent reasoning\. This separation allows long\-context history to be reused across requests while keeping each online reasoning context focused and grounded in the original agent interaction\.
## 4Experiments
### 4\.1Evaluation on the LoCoMo Benchmark Dataset
We report evaluation\-model accuracy using the Memora\-compatible protocol\([Xia et al\., 2026](https://arxiv.org/html/2609.27279#bib.bib19)\)\. Because answers are open\-ended, BLEU and F1 which measure lexical overlap rather than semantic correctness and do not represent a convincing measure; we therefore use an LLM\-based, fixed binary evaluation protocol\. LLM model\-based evaluation is a scalable proxy for human assessment, while known evaluator biases motivate reporting two evaluation models separately\([Zheng et al\., 2023](https://arxiv.org/html/2609.27279#bib.bib17)\)\. The model used for answer generation in both offline and online stages is GPT\-4\.1\-mini\. We report the GPT\-4o\-mini and Qwen3\-32B evaluation\-model results separately rather than averaging them\. Benchmark definitions and category mappings are provided in Appendix[A](https://arxiv.org/html/2609.27279#A1)\.
EnSIMem is strongest on single\-hop and temporal questions, while open\-domain questions remain the most challenging slice\. Overall, GPT\-4o\-mini evaluates the answer and gives an overall accuracy of 90\.6%, and Qwen3\-32B gives an overall accuracy of 90\.0%\. We report these evaluation\-model results separately; the GPT\-4o\-mini row remains the directly comparable result against the published Memora numbers\.
Table 1:Evaluation: LLM\-model\-based accuracy on LoCoMo\. Published baseline and MEMORA scores are transcribed from the Memora benchmark table; asterisks mark baselines reported there from prior work\. We report the same EnSIMem answers under GPT\-4o\-mini and Qwen3\-32B evaluation models\.
### 4\.2Evaluation on the LongMemEval Benchmark Dataset
We report evaluation\-model accuracy using the same two\-model protocol\. Comparison\-row provenance is provided in Appendix[A](https://arxiv.org/html/2609.27279#A1)\.
Table 2:Evaluation: LLM\-model\-based accuracy on LongMemEval\. “SS” denotes single\-session\. Published comparison rows are transcribed from Memora; EnSIMem reports the results from both evaluation models separately\. Bold denotes the highest value in each column; underlining denotes the second\-highest value\.GPT\-4o\-mini reaches 92\.8% overall on LongMemEval, improving over MEMORA \(P\) by 5\.41 points; Qwen3\-32B also reaches 92\.8%, improving by 5\.44 points\. The largest gains are on SS\-assistant and multi\-session \(\+13\.1/\+14\.8 points for GPT\-4o\-mini and \+16\.0/\+13\.2 for Qwen3\-32B\), while the remaining categories differ by at most 1\.30 points except SS\-pref\., which is 3\.4 points higher\. We therefore report the two evaluation\-model results separately rather than averaging them\.
### 4\.3Efficiency
We instrument Steps 4–6 on the current LoCoMo run, recording wall\-clock latency, context\-token counts, and retrieval search steps\. The instrumentation does not change ranking, stopping rules, or generated answers\.
Table 3:Online efficiency of EnSIMem on LoCoMo\.The system averages 1\.92 retrieval steps per query\. Planning accounts for 71\.1% of mean end\-to\-end latency, compared with 11\.0% for retrieval and 17\.9% for generation\. This overhead constructs a more complete short reasoning buffer by identifying the property granularity, reasoning type, and evidence budget\. Memora reports 5\.70 seconds mean end\-to\-end latency and 4\.61 seconds mean search latency\([Xia et al\., 2026](https://arxiv.org/html/2609.27279#bib.bib19)\)\. However, this cannot be treated as a direct comparison with our EnSIMem, as we use different processing engines and execution settings\.
### 4\.4Ablation Study
We conduct controlled LoCoMo ablations, changing one design choice while fixing the answer model, evaluation model, and evaluation protocol\. Each comparison isolates one source of the final gain rather than retuning the full pipeline\.
#### Episode granularity\.
We compare theme\-coherent episodes with per\-turn and per\-session units\. This tests whether coherent middle\-sized units improve the relevance–noise trade\-off for an agent’s reasoning buffer\. The results are shown in Figure[2](https://arxiv.org/html/2609.27279#S4.F2)\. EnSIMem obtains 96\.59% accuracy, compared with 94\.32% for per\-session episodes and 90\.91% for per\-turn episodes\. Thus, theme\-coherent episodes improve accuracy by 2\.27 and 5\.68 percentage points over the two alternatives, respectively\. This isolates memory\-unit granularity: per\-turn units can lose local references, whereas per\-session units can admit unrelated context\.

Figure 2:Ablation results on LoCoMo\. The left panel compares episode granularity, showing that theme\-coherent episodes outperform per\-turn and per\-session units\. The right panel compares four property granularities: Fine, Broad \(EnSIMem setting\), Broad\+, and Very broad\. Broad\+ is a slightly broader variant that merges additional action\-like predicates into the shared property*activity*, while still preserving intermediate property distinctions\. Higher accuracy is better\.
#### Property granularity\.
Using the same episodes, we compare four extraction granularities\. Fine preserves surface predicates; Broad\+ is slightly broader than EnSIMem’s Broad prompt and merges additional action\-like predicates into*activity*; Very broad maps more properties to*fact*or broad concepts\. Accuracy is 92\.05%, 96\.59%, 95\.45%, and 90\.91% for Fine, Broad \(EnSIMem\), Broad\+, and Very broad, respectively, showing that moderate normalization helps while excessive coarsening hurts retrieval\. These results suggest that moderate normalization improves property alignment, while overly coarse properties remove distinctions needed for retrieval\.
#### Structured matching versus dense retrieval\.
Dense retrieval ranks complete episodes by semantic similarity, whereas EnSIMem first matches entity, type, and property fields and links records to source episodes\. With the answer and evaluation models fixed, EnSIMem reaches 96\.59% versus 86\.36% for dense retrieval \(Table[4](https://arxiv.org/html/2609.27279#S4.T4)\), showing the benefit of the structured access path\. It therefore tests whether explicit index structure contributes beyond dense semantic similarity alone\.
Table 4:Retrieval\-route ablation on LoCoMo\. Both variants use the same answer model and GPT\-4o\-mini evaluation model\.Figure 3:Fixed top\-kkbudgets versus EnSIMem’s adaptive budget on LoCoMo\. The dashed line shows the adaptive result\. Higher accuracy is better\.
#### Retrieval Episode Strategy\.
We compare the adaptive budget with fixed top\-kkpolicies fork∈\{1,3,5,8,15,20\}k\\in\\\{1,3,5,8,15,20\\\}\. Fixed retrieval rises from 52\.27% atk=1k=1to 95\.45% atk=8k=8before saturating or declining, while EnSIMem reaches 96\.59% \(Figure[3](https://arxiv.org/html/2609.27279#S4.F3)\)\. This isolates query\-type\-aware allocation: fixed budgets trade recall and noise uniformly, whereas the adaptive policy can spend evidence where the question requires coverage\.
## 5Discussion, Limitations, and Future Work
Our study focuses on episodic conversational memory and does not yet cover procedural or semantic memory\. The pipeline relies on LLM\-based episode segmentation, property extraction, and query planning; although our ablations show robust trends, different models or prompts may produce different intermediate representations, motivating future work on confidence estimation and cross\-model validation\.
LoCoMo and LongMemEval primarily evaluate conversational recall and reasoning, rather than skill execution or persistent factual updates\. In addition, evaluation\-model preferences may affect absolute scores, while runtime comparisons depend on the underlying processing engine and deployment configuration\.
A natural extension is to integrate EnSIMem with procedural and semantic memory\. Procedural records could represent reusable skills with goals, preconditions, actions, outcomes, and provenance, while semantic records could maintain entity–property facts with validity intervals and supporting episodes\. An agent harness could coordinate these stores through shared entity identifiers and provenance links, selecting the appropriate memory type for each request\. Future evaluation should therefore include skill reuse, fact consolidation, temporal updates, and cross\-memory reasoning\.
## 6Conclusion
EnSIMem reframes long\-term conversational memory as a short\-context reasoning problem\. Theme\-coherent episodes organize local context, while entity–property records provide precise handles that connect each request to the relevant source evidence\. Requirement\-aware planning and adaptive retrieval select only the evidence needed for point, temporal, compositional, and aggregation questions\. Unlike generalized summaries or anonymous chunks, the reasoning buffer retains the original text, provenance, and available multimodal evidence, reducing irrelevant context while preserving the details required for grounded answers\. This evidence\-preserving design provides a foundation for extending structured memory access beyond episodic memory\.
## AI Use Disclosure
In this work, we used generative AI tools to polish the writing, verify grammar errors, and implement and debug software\. We also used GPT\-4\.1\-mini as the answer\-generation model in our benchmark experiments and GPT\-4o\-mini and Qwen3\-32B as evaluation models for assessing generated answers\. These model\-based evaluations were conducted as part of the experimental protocol and were not used to create synthetic training data or alter the ground\-truth annotations\.
We did not use generative AI tools to generate synthetic datasets, provide proofs, or establish mathematical claims without author verification\. All AI\-assisted code was reviewed, executed, and tested by the authors\. The authors independently checked the manuscript text, citations, mathematical notation, figures, experimental configurations, and reported results\. We take full responsibility for the final content of this paper\.
## References
- Asaiet al\.\(2024\)A\. Asai, Z\. Wu, Y\. Wang, A\. Sil, and H\. HajishirziSelf\-RAG: learning to retrieve, generate, and critique through self\-reflection\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2609.27279#S1.p3.1)\.
- Baiet al\.\(2025\)Y\. Bai, S\. Tu, J\. Zhang, H\. Peng, X\. Wang, X\. Lv, S\. Cao, J\. Xu, L\. Hou, Y\. Dong, J\. Tang, and J\. LiLongBench v2: towards deeper understanding and reasoning on realistic long\-context multitasks\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),External Links:[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.183)Cited by:[§1](https://arxiv.org/html/2609.27279#S1.p3.1)\.
- Chhikaraet al\.\(2025\)P\. Chhikara, D\. Khant, S\. Aryan, T\. Singh, and D\. YadavMem0: building production\-ready ai agents with scalable long\-term memory\.arXiv preprint arXiv:2504\.19413\.Cited by:[§A\.2](https://arxiv.org/html/2609.27279#A1.SS2.p2.1.6.3.1.1),[§2](https://arxiv.org/html/2609.27279#S2.SS0.SSS0.Px1.p1.1)\.
- Duet al\.\(2025\)Y\. Du, M\. Tian, S\. Ronanki, S\. Rongali, S\. Bodapati, A\. Galstyan, A\. Wells, R\. Schwartz, E\. A\. Huerta, and H\. PengContext length alone hurts llm performance despite perfect retrieval\.arXiv preprint arXiv:2510\.05381\.Cited by:[§1](https://arxiv.org/html/2609.27279#S1.p3.1)\.
- Edgeet al\.\(2024\)D\. Edge, H\. Trinh, N\. Cheng, J\. Bradley, A\. Chao, A\. Mody, S\. Truitt, D\. Metropolitansky, R\. O\. Ness, and J\. LarsonFrom local to global: a graph RAG approach to query\-focused summarization\.arXiv preprint arXiv:2404\.16130\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2404.16130)Cited by:[§1](https://arxiv.org/html/2609.27279#S1.p3.1),[§2](https://arxiv.org/html/2609.27279#S2.SS0.SSS0.Px2.p1.1)\.
- Heet al\.\(2024\)X\. He, Y\. Tian, Y\. Sun, N\. V\. Chawla, T\. Laurent, Y\. LeCun, X\. Bresson, and B\. HooiG\-Retriever: retrieval\-augmented generation for textual graph understanding and question answering\.arXiv preprint arXiv:2402\.07630\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2402.07630)Cited by:[§1](https://arxiv.org/html/2609.27279#S1.p6.1)\.
- Huet al\.\(2024\)Y\. Hu, Z\. Lei, Z\. Zhang, B\. Pan, C\. Ling, and L\. ZhaoGRAG: graph retrieval\-augmented generation\.arXiv preprint arXiv:2405\.16506\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2405.16506)Cited by:[§1](https://arxiv.org/html/2609.27279#S1.p6.1)\.
- Izacardet al\.\(2022\)G\. Izacard, M\. Caron, L\. Hosseini, S\. Riedel, P\. Bojanowski, A\. Joulin, and E\. GraveUnsupervised dense information retrieval with contrastive learning\.Transactions on Machine Learning Research\.Cited by:[§1](https://arxiv.org/html/2609.27279#S1.p3.1)\.
- Jiménez Gutiérrezet al\.\(2024\)B\. Jiménez Gutiérrez, Y\. Shu, Y\. Gu, M\. Yasunaga, and Y\. SuHippoRAG: neurobiologically inspired long\-term memory for large language models\.arXiv preprint arXiv:2405\.14831\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2405.14831)Cited by:[§A\.2](https://arxiv.org/html/2609.27279#A1.SS2.p2.1.4.3.1.1),[§2](https://arxiv.org/html/2609.27279#S2.SS0.SSS0.Px2.p1.1)\.
- Kanget al\.\(2025\)J\. Kang, M\. Ji, Z\. Zhao, and T\. BaiMemory OS of AI agent\.arXiv preprint arXiv:2506\.06326\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2506.06326)Cited by:[§1](https://arxiv.org/html/2609.27279#S1.p1.1),[§2](https://arxiv.org/html/2609.27279#S2.SS0.SSS0.Px1.p1.1)\.
- Karpukhinet al\.\(2020\)V\. Karpukhin, B\. Oguz, S\. Min, P\. Lewis, L\. Wu, S\. Edunov, D\. Chen, and W\. YihDense passage retrieval for open\-domain question answering\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing,pp\. 6769–6781\.External Links:[Document](https://dx.doi.org/10.18653/v1/2020.emnlp-main.550)Cited by:[§A\.2](https://arxiv.org/html/2609.27279#A1.SS2.p2.1.3.3.1.1),[§1](https://arxiv.org/html/2609.27279#S1.p3.1),[§2](https://arxiv.org/html/2609.27279#S2.SS0.SSS0.Px2.p1.1)\.
- Khattab and Zaharia \(2020\)O\. Khattab and M\. ZahariaColBERT: efficient and effective passage search via contextualized late interaction over BERT\.InProceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval,pp\. 39–48\.External Links:[Document](https://dx.doi.org/10.1145/3397271.3401075)Cited by:[§1](https://arxiv.org/html/2609.27279#S1.p3.1),[§2](https://arxiv.org/html/2609.27279#S2.SS0.SSS0.Px2.p1.1)\.
- Lewiset al\.\(2020\)P\. Lewis, E\. Perez, A\. Piktus, F\. Petroni, V\. Karpukhin, N\. Goyal, H\. Küttler, M\. Lewis, W\. Yih, T\. Rocktäschel, S\. Riedel, and D\. KielaRetrieval\-augmented generation for knowledge\-intensive NLP tasks\.InAdvances in Neural Information Processing Systems,Vol\.33,pp\. 9459–9474\.Cited by:[§A\.2](https://arxiv.org/html/2609.27279#A1.SS2.p2.1.3.3.1.1),[§1](https://arxiv.org/html/2609.27279#S1.p3.1),[§2](https://arxiv.org/html/2609.27279#S2.SS0.SSS0.Px2.p1.1)\.
- Liuet al\.\(2024\)N\. F\. Liu, K\. Lin, J\. Hewitt, A\. Paranjape, M\. Bevilacqua, F\. Petroni, and P\. LiangLost in the middle: how language models use long contexts\.Transactions of the Association for Computational Linguistics12,pp\. 157–173\.External Links:[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00638)Cited by:[§1](https://arxiv.org/html/2609.27279#S1.p3.1)\.
- Maharanaet al\.\(2024\)A\. Maharana, D\. Lee, S\. Tulyakov, M\. Bansal, F\. Barbieri, and Y\. FangEvaluating very long\-term conversational memory of LLM agents\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 13829–13849\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.747)Cited by:[§A\.1](https://arxiv.org/html/2609.27279#A1.SS1.p2.1.2.3.1.1),[§1](https://arxiv.org/html/2609.27279#S1.p2.1)\.
- Menget al\.\(2026\)X\. Meng, J\. Sun, J\. R\. Parekh, and J\. HanEnSI\-RAG: entity\-structure\-indexed retrieval\-augmented generation for long\-document question answering\.arXiv preprint arXiv:2608\.21252\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2608.21252)Cited by:[§1](https://arxiv.org/html/2609.27279#S1.p4.1),[§2](https://arxiv.org/html/2609.27279#S2.SS0.SSS0.Px2.p1.1)\.
- Packeret al\.\(2023\)C\. Packer, S\. Wooders, K\. Lin, V\. Fang, S\. G\. Patil, I\. Stoica, and J\. E\. GonzalezMemGPT: towards LLMs as operating systems\.arXiv preprint arXiv:2310\.08560\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2310.08560)Cited by:[§1](https://arxiv.org/html/2609.27279#S1.p1.1),[§2](https://arxiv.org/html/2609.27279#S2.SS0.SSS0.Px1.p1.1)\.
- Parekhet al\.\(2025\)J\. R\. Parekh, P\. Jiang, and J\. HanStructure\-augmented reasoning generation\.arXiv preprint arXiv:2506\.08364\.Cited by:[§1](https://arxiv.org/html/2609.27279#S1.p4.1)\.
- Parket al\.\(2023\)J\. S\. Park, J\. C\. O’Brien, C\. J\. Cai, M\. R\. Morris, P\. Liang, and M\. S\. BernsteinGenerative agents: interactive simulacra of human behavior\.InProceedings of the 36th Annual ACM Symposium on User Interface Software and Technology,External Links:[Document](https://dx.doi.org/10.1145/3586183.3606763)Cited by:[§1](https://arxiv.org/html/2609.27279#S1.p1.1),[§2](https://arxiv.org/html/2609.27279#S2.SS0.SSS0.Px1.p1.1)\.
- Rasmussenet al\.\(2025\)P\. Rasmussen, P\. Paliychuk, T\. Beauvais, J\. Ryan, and D\. ChalefZep: a temporal knowledge graph architecture for agent memory\.arXiv preprint arXiv:2501\.13956\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2501.13956)Cited by:[§A\.2](https://arxiv.org/html/2609.27279#A1.SS2.p2.1.5.3.1.1),[§2](https://arxiv.org/html/2609.27279#S2.SS0.SSS0.Px1.p1.1)\.
- Shankaret al\.\(2024\)S\. Shankar, T\. Chambers, T\. Shah, A\. G\. Parameswaran, and E\. WuDocETL: agentic query rewriting and evaluation for complex document processing\.arXiv preprint arXiv:2410\.12189\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2410.12189)Cited by:[§1](https://arxiv.org/html/2609.27279#S1.p5.1)\.
- Shinnet al\.\(2023\)N\. Shinn, F\. Cassano, A\. Gopinath, K\. Narasimhan, and S\. YaoReflexion: language agents with verbal reinforcement learning\.InAdvances in Neural Information Processing Systems,Cited by:[§2](https://arxiv.org/html/2609.27279#S2.SS0.SSS0.Px1.p1.1)\.
- Trivediet al\.\(2023\)H\. Trivedi, N\. Balasubramanian, T\. Khot, and A\. SabharwalInterleaving retrieval with chain\-of\-thought reasoning for knowledge\-intensive multi\-step questions\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 10014–10037\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.acl-long.557)Cited by:[§1](https://arxiv.org/html/2609.27279#S1.p3.1),[§1](https://arxiv.org/html/2609.27279#S1.p5.1)\.
- Wanget al\.\(2023\)G\. Wang, Y\. Xie, Y\. Jiang, A\. Mandlekar, C\. Xiao, Y\. Zhu, L\. Fan, and A\. AnandkumarVoyager: an open\-ended embodied agent with large language models\.arXiv preprint arXiv:2305\.16291\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2305.16291)Cited by:[§2](https://arxiv.org/html/2609.27279#S2.SS0.SSS0.Px1.p1.1)\.
- Wanget al\.\(2024\)M\. Wang, L\. Chen, C\. Fu, S\. Liao, X\. Zhang, B\. Wu, H\. Yu, N\. Xu, L\. Zhang, R\. Luo, Y\. Li, M\. Yang, F\. Huang, and Y\. LiLeave no document behind: benchmarking long\-context LLMs with extended multi\-document QA\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,External Links:[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.322)Cited by:[§1](https://arxiv.org/html/2609.27279#S1.p3.1)\.
- Wuet al\.\(2024\)D\. Wu, H\. Wang, W\. Yu, Y\. Zhang, K\. Chang, and D\. YuLongMemEval: benchmarking chat assistants on long\-term interactive memory\.arXiv preprint arXiv:2410\.10813\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2410.10813)Cited by:[§A\.1](https://arxiv.org/html/2609.27279#A1.SS1.p2.1.3.3.1.1),[§1](https://arxiv.org/html/2609.27279#S1.p2.1)\.
- Xiaet al\.\(2026\)M\. Xia, X\. Zhang, S\. Dixit, P\. Harimurugan, R\. Wang, V\. Ruhle, R\. Sim, C\. Bansal, and S\. RajmohanMemora: a harmonic memory representation balancing abstraction and specificity\.arXiv preprint arXiv:2602\.03315\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2602.03315)Cited by:[§A\.1](https://arxiv.org/html/2609.27279#A1.SS1.p2.1.2.3.1.1),[§A\.1](https://arxiv.org/html/2609.27279#A1.SS1.p2.1.3.3.1.1),[§A\.2](https://arxiv.org/html/2609.27279#A1.SS2.p2.1.7.3.1.1),[§A\.2](https://arxiv.org/html/2609.27279#A1.SS2.p2.1.8.3.1.1),[§A\.2](https://arxiv.org/html/2609.27279#A1.SS2.p2.1.9.3.1.1),[Appendix C](https://arxiv.org/html/2609.27279#A3.p1.1),[§2](https://arxiv.org/html/2609.27279#S2.SS0.SSS0.Px3.p1.1),[§4\.1](https://arxiv.org/html/2609.27279#S4.SS1.p1.1),[§4\.3](https://arxiv.org/html/2609.27279#S4.SS3.p2.1)\.
- Xuet al\.\(2025\)W\. Xu, Z\. Liang, K\. Mei, H\. Gao, J\. Tan, and Y\. ZhangA\-mem: agentic memory for LLM agents\.arXiv preprint arXiv:2502\.12110\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2502.12110)Cited by:[§1](https://arxiv.org/html/2609.27279#S1.p1.1),[§2](https://arxiv.org/html/2609.27279#S2.SS0.SSS0.Px1.p1.1)\.
- Zhenget al\.\(2023\)L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. P\. Xing, H\. Zhang, J\. E\. Gonzalez, and I\. StoicaJudging llm\-as\-a\-judge with mt\-bench and chatbot arena\.InAdvances in Neural Information Processing Systems,External Links:[Link](https://arxiv.org/abs/2306.05685)Cited by:[§4\.1](https://arxiv.org/html/2609.27279#S4.SS1.p1.1)\.
- Zhonget al\.\(2023\)W\. Zhong, L\. Guo, Q\. Gao, H\. Ye, and Y\. WangMemoryBank: enhancing large language models with long\-term memory\.arXiv preprint arXiv:2305\.10250\.Cited by:[§2](https://arxiv.org/html/2609.27279#S2.SS0.SSS0.Px1.p1.1)\.
## Appendix AAppendix
### A\.1Benchmark overview
Table 5\.Benchmark scope, question families, and evaluation protocol used in this paper\.
### A\.2Baseline and method overview
Table 6\.Baseline and method comparison used in the benchmark tables\.
## Appendix BCase Studies
We present three representative LoCoMo cases in which the evaluation model labels EnSIMem correct while labeling Memora incorrect\. The cases are selected to expose three distinct mechanisms: cross\-episode evidence coverage, temporal grounding, and fine\-grained entity\-property localization\. Each table records the question, the required operation, the memory representation, the retrieved evidence, the generated responses, and the specific failure mode of the competing system\.
Table 7\.Case study 1: cross\-episode evidence coverage for a multi\-fact question\.
Table 8\.Case study 2: temporal grounding from a relative\-time expression\.
Table 9\.Case study 3: entity\-property specificity for multimodal art evidence\.
## Appendix CEvaluation Model Rationale
Our evaluation follows the answer\-scoring protocol used by Memora and its underlying Mem0 evaluation implementation\([Xia et al\., 2026](https://arxiv.org/html/2609.27279#bib.bib19)\)\. This choice is appropriate for long\-term agent memory because the benchmark answers are open\-ended: a correct answer may paraphrase the reference, use a relative temporal expression, or include additional explanation without matching the reference string word for word\.
#### LoCoMo decision rule\.
For every query, the evaluation program passes the question, the gold answer, and the generated answer to an evaluation model\. The Memora/Mem0\-compatible prompt asks whether the generated answer is semantically consistent with the gold answer and requests a binaryCORRECTorWRONGlabel in JSON\. The implementation mapsCORRECTto one andWRONGto zero, then aggregates these decisions overall and by LoCoMo category\. Category 5 is excluded according to the Memora\-compatible scoring policy\. Thus, the reported LoCoMo number measures answer\-level semantic correctness rather than token overlap\.
#### LongMemEval decision rule\.
LongMemEval uses the corresponding Memora\-style task\-specific checker\. The evaluation program again separates answer generation from evaluation, but selects a prompt according to the question type: preference questions use a personalization rubric, knowledge\-update questions check the updated answer, temporal questions include their temporal tolerance, and unanswerable items check whether the model correctly abstains\. The evaluator receives the question, reference answer or rubric, and generated response, and returnsyesorno; the implementation mapsyesto correct and aggregates accuracy by question type and overall\. This preserves the benchmark’s intended semantics while keeping the evaluation procedure independent of the retrieval pipeline\.
#### Why this is preferable to lexical metrics\.
BLEU and F1 primarily reward shared surface tokens\. They can penalize a valid paraphrase, fail to recognize equivalent date formats, and provide little evidence that the answer identifies the requested entity, property, or event\. The evaluation\-model prompt instead exposes the complete question–gold–prediction triple and explicitly allows semantically equivalent wording and time expressions\. This makes the protocol better aligned with the actual objective of an agent\-memory system: recovering the right fact or event for the user’s request\.
#### Reproducibility and model dependence\.
LoCoMo evaluation uses temperature zero, a fixed seed of 42, and a fixed Memora\-compatible prompt; the LongMemEval checker uses temperature zero and its fixed task\-specific prompts\. The same protocol and aggregation code are used across methods within each benchmark\. Because model\-based evaluation can still reflect evaluator\-specific preferences, we report GPT\-4o\-mini and Qwen3\-32B results separately rather than treating their arithmetic mean as a new ground truth\. Agreement in the qualitative trends across the two evaluation models provides a more transparent robustness check while preserving the identity of each evaluation model\.
#### Relation to the benchmark comparison\.
Using the Memora\-compatible protocol keeps our LoCoMo comparison on the same answer\-level scale as the published Memora rows\. It also makes the source of every score explicit: the answer model generates the response, the evaluation model makes the binary semantic decision, and the reported accuracy is the fraction of correct decisions\. This separation avoids conflating answer generation with evaluation and allows future work to replace or add evaluation models without changing the stored answers or retrieval pipeline\.
## Appendix DEvaluation\-Model Consistency
We additionally examine whether the reported scores are stable across the two evaluation models\. For each benchmark category, letgcg\_\{c\}andqcq\_\{c\}denote the accuracy obtained with GPT\-4o\-mini and Qwen3\-32B, respectively, and define the signed differenceΔc=qc−gc\\Delta\_\{c\}=q\_\{c\}\-g\_\{c\}in percentage points \(pp\)\. We report the mean signed difference, mean absolute difference \(MAE\), the population standard deviation of the category differences, the largest absolute category difference, the gap between the two overall scores, and the Pearson correlation across category accuracies\. The analysis uses the unrounded category\-level results underlying the main tables before their three\-significant\-digit display; it does not treat the arithmetic mean of the two evaluation models as a new score\.
The two evaluation models therefore give highly similar aggregate conclusions\. On LoCoMo, the overall scores differ by only0\.6100\.610pp \(90\.6% versus 90\.0%\), and the category\-level MAE is2\.112\.11pp\. The larger LoCoMo discrepancy is concentrated in multi\-hop questions, where Qwen3\-32B is5\.465\.46pp lower; the remaining category gaps are at most2\.092\.09pp\. On LongMemEval, the overall gap is only0\.0300\.030pp \(92\.8% versus 92\.8% in three significant figures\), with a smaller category\-level MAE of1\.371\.37pp\. The largest LongMemEval difference is3\.373\.37pp on single\-session preference questions, while temporal reasoning, knowledge update, and single\-session user questions differ by at most0\.2500\.250pp\. The high category\-level correlations \(r=0\.946r=0\.946for LoCoMo andr=0\.936r=0\.936for LongMemEval\), together with the small absolute gaps, indicate that the headline findings are consistent across the two evaluation models even though individual categories can show model\-specific sensitivity\.
Table 10:Consistency of GPT\-4o\-mini and Qwen3\-32B evaluation\-model scores\.Δ\\Deltais defined as Qwen3\-32B minus GPT\-4o\-mini; all difference columns are in percentage points\. The correlation is computed across benchmark categories, excluding the overall column\.This is an aggregate stability analysis rather than a claim of perfect per\-query agreement\. A paired reliability analysis would require retaining both binary labels for every identical query\. With those labels, one could additionally report exact agreement, Cohen’sκ\\kappa, a paired McNemar test, and bootstrap confidence intervals\. We do not infer those quantities from category totals alone; instead, we report the two evaluation\-model results separately and use the agreement of their aggregate trends as the robustness check\.
## Appendix EPrompts
This appendix records the core prompts used by EnSIMem\. The text inside each box is the prompt template passed to the corresponding LLM stage; braces denote runtime substitutions such as a conversation, an episode, a query, or a retrieval plan\. We show the prompts that determine the memory representation, query decomposition, evidence\-grounded answering, and evaluation\. Repair and validation prompts are invoked only when a model returns malformed JSON and follow the same schemas shown here\. Structured matching, dense fallback, evidence\-budget expansion, and score aggregation are deterministic operations and do not invoke an LLM\. The corresponding source files areLoCoMo/prompts\.py,LoCoMo/08\_memora\_llm\_judge\.py, andLongMemEval/memora\_evaluation\.py; the runtime substitutions shown below are made by the pipeline scripts\.
### E\.1Offline memory construction
#### Theme\-coherent episode partitioning\.
The partition prompt creates contiguous evidence units before any entity\-property records are extracted\. It preserves dialogue order and keeps image metadata attached to the turn that introduced it\.
Youpartitionacompletedialoguesessionintocontiguoustheme\-coherentmemoryepisodes\.Theseboundariesdefineimmutableevidenceunits\.ReturnoneJSONobjectonly\.
Partitionthissession\.
Amemoryepisodeisonecontiguousspanmainlyfocusedononecoherententity,event,goal,orboundedtopic\.Usetheme\_typeentity,event,ortopic\.Splitonlyonagenuinesemanticshift;keepgreetings,acknowledgements,clarificationquestions,andshortfollow\-upswiththematerialtheysupport\.
Structuralrules:
\-Covereverysupplieddialogueturnexactlyonceandinorder,withoutgapsoroverlaps\.
\-Donotsplitmerelybecausethespeakerchanges\.
\-Avoidone\-turnfragmentsunlessaturnclearlyintroducesaseparatedurabletheme\.
\-Keeptext,imagecaptions,image\-retrievaldescriptions,andimageURLsattachedtotheirDIA\.
\-Copystart\_dia\_idandend\_dia\_idexactly\.
ReturnJSONonly:
\{”segments”:\[\{”start\_dia\_id”:”D1:1”,”end\_dia\_id”:”D1:5”,”theme”:”concisecanonicaltheme”,”theme\_type”:”entity\|event\|topic”,”theme\_description”:”one\-sentencescope”,”boundary\_reason”:”session\_startorsemanticshift”\}\]\}
Conversation:\{conversation\_id\}
Session:\{session\_id\}
Observedat:\{observed\_at\}
Turns:
\{session\_text\}
#### Entity\-property index extraction\.
This prompt is query\-independent: it extracts atomic records once and keeps the source episode as the only answer evidence\. The property policy is deliberately intermediate\-grained, while values and source DIA identifiers retain concrete detail\.
Youarethehigh\-recall,precision\-preservinginformation\-extractionstageofanepisodicmemorysystem\.Buildaquery\-independententity\-structuredindexfromonecompletethemeepisode\.Theepisoderemainstheonlyanswerevidence;recordsarenavigationhandles\.Followtheminimum\-sufficientpredicatepolicyandreturnonevalidJSONobjectonly\.
Extractanexhaustivesetofatomicrecordsfromtheepisodebelow\.
Readeverydialogueturnandeverydisplayedmetadatafield\.Imagecaptionsandimage\-retrievaldescriptionsareauxiliarytextualevidence;animageURLisopaqueprovenanceandmustneverbeusedforoutsidelookup\.
Extractionrequirements:
\-ScaneachDIAindependentlyandemitarecordforeveryexplicitfactualclause\.
\-ResolveI/my/meusingthespeakeronthatturnandciteexactevidence\_dia\_ids\.
\-Preservethespeaker\-to\-entityrelationforexplicitfirst\-personactionsandemitexplicitinverserelationsonlywhenentailed\.
\-Useoneentityandoneminimum\-sufficientreusableproperty\.Keepdistinctivepredicatessuchasresearch,travel,attend,read,paint,checkup,andsupport;donotcollapsethemallintoactivityorconcatenatethesubjectintotheproperty\.
\-Putconcretesourcedetailinvalue\.Preserverelativetime,modality,polarity,conditions,andimagemetadata\.
\-Donotinferfactsfromworldknowledge,episodeadjacency,oranimageURL\.
Returnoneobjectonly:
\{”records”:\[\{”entity”:”…”,”entity\_type”:”person\|organization\|place\|object\|event\|activity\|concept\|other”,”property”:”shortpredicate”,”value”:”explicitvalueorempty”,”condition\_property”:”…orempty”,”condition\_value”:”…orempty”,”property\_kind”:”aspect\|relation”,”modality”:”observed\|planned\|desired\|hypothetical\|negated\|uncertain”,”valid\_time”:”…orempty”,”source”:”speaker”,”evidence\_dia\_ids”:\[”D1:1”\],”projection\_kind”:”direct\|inverse\|state\_projection\|event\_access”,”confidence”:0\.0\}\]\}
Episodemetadata:
EpisodeID:\{episode\_id\}
Conversation:\{conversation\_id\}
Session:\{session\_id\}
Observedat:\{observed\_at\}
Episodetheme:\{episode\_theme\}\(\{episode\_theme\_type\}\)
Episodeturns:
\{episode\_text\}
### E\.2Online retrieval and response generation
#### Requirement\-aware query planning\.
The planner converts a question into explicit requirements and an evidence\-seeking hop graph\. It uses the same schema as the index, marks aggregation questions asall\_matching, and leaves unknown answer values empty\.
Youaretheexhaustivequery\-decompositionstageofanepisodicmemorysystem\.Convertthequestionintosearchablerequirementsandanevidence\-seekinghopgraphusingthesame\[entity\]\[entity\_type\]\[property:value\]\[condition:value\]schemaasthememoryindex\.Preserveeverymeaningfuldiscriminator,butneverguessananswer\.Donotanswerthequestion\.ReturnexactlyonevalidJSONobject\.
Buildanexhaustiveretrievalplanforthequestionbelow\.
Planningrules:
\-Includethenamedentity,therelationshipthatdisambiguatesthetarget,therequestedproperty,andallexplicittime,quantity,andlocationconstraints\.
\-Useminimum\-sufficientpropertiesalignedwiththeindex;donotinventsubject\-prefixedorcompoundlabels\.
\-Keepunknownanswervaluesempty\.Useonlyinformationstatedinthequestionoradeclaredbridge\.
\-Addseparatehopsforindependentlyusefuldiscriminators\.Forexample,adaughter’sbirthdayrequiresrelationship=daughter,property=birthday,andproperty=time\.
\-Usereasoning\_type=aggregationorcomparisonandretrieval\_scope=all\_matchingforcounts,exhaustivelists,frequencies,intervals,andfirst/secondquestions\.Otherwiseusepointretrieval\.
\-Preserverelativetimeandexplicitconditions\.Donotanswerthequestioninthisstage\.
Returnoneobjectonly:
\{”answer\_target”:\{”type”:”time\|value\|entity\|location\|count\|list\|boolean\|likelihood\|explanation”,”description”:”…”\},”reasoning\_type”:”direct\|aggregation\|comparison\|temporal\|causal\|counterfactual\|commonsense\_inference\|multi\_hop\|unanswerable”,”retrieval\_scope”:”point\|all\_matching”,”required\_properties”:\[\{”entity”:”…”,”entity\_type”:”…”,”broad\_property”:”…”,”property\_text”:”…”,”value”:”…”,”role”:”entity\|relation\|answer\_property\|time\|constraint”\}\],”hops”:\[\{”hop\_id”:”h1”,”purpose”:”…”,”anchor”:\{”entity”:””,”entity\_type”:””,”property”:””,”value”:””,”condition\_property”:””,”condition\_value”:””\},”depends\_on”:\[\],”bridge\_request”:””\}\]\}
Observedindexpropertyvocabulary:
\{index\_property\_vocabulary\}
Question:\{question\}
#### Evidence\-grounded answer generation\.
The answer model receives complete retrieved episodes, not only the structured records\. The structured records and hop plan are navigation aids; the final answer must be supported by the original dialogue and attached image metadata\.
Answeronlyfromthesuppliedcompleteoriginalconversationepisodes\.Thestructuredindexesandhopplanarenavigationaids,notevidence\.Readevidencefromeveryhop,combinefactswhentheplanismulti\-hop,andkeepentitiescorrectlybound\.Donotuseoutsideknowledge\.TreatcaptionsandretrievaldescriptionsastextualevidenceattachedtotheirexactDIA;treatimageURLsasopaqueprovenance\.
Forlist,count,frequency,interval,andfirst/secondquestions,inventoryeveryqualifyingeventacrossallsuppliedepisodesbeforeanswering\.Fortemporalquestions,bindtherequestedeventfirstandpreserveapproximatesourcewording\.Iftheevidencedoesnotestablishtheanswer,outputexactly:Unknown\.Returnonlytheshortestsufficientanswerwithoutexplanation\.
Question:\{question\}
Retrievalplan:
\{plan\}
Completeoriginalepisodesretrievedacrossallhops:
\{episodes\}
Answer:
### E\.3Evaluation prompts
The benchmark evaluation is separate from answer generation\. The evaluator receives the question, the benchmark gold answer \(or rubric\), and the generated answer\. It does not receive the retrieval plan or hidden retrieval diagnostics\. LoCoMo uses the Memora\-compatible binary prompt below; LongMemEval selects the corresponding task\-specific variant according to the question type\.
YourtaskistolabelananswertoaquestionasCORRECTorWRONG\.Youwillbegiven\(1\)aquestion,\(2\)agoldanswer,and\(3\)ageneratedanswer\.Begenerouswithgrading:alongeransweriscorrectwhenittouchesthesametopicasthegoldanswer,andatimeansweriscorrectwhenitreferstothesamedateorperiodeveniftheformatdiffers\.
Question:\{question\}
Goldanswer:\{gold\_answer\}
Generatedanswer:\{generated\_answer\}
ReturnonlyaJSONobjectwiththekey”label”andthevalueCORRECTorWRONG\.
Iwillgiveyouaquestion,areferenceanswerorrubric,andamodelresponse\.Answeryesiftheresponseissemanticallycorrectandnootherwise\.Theresponsemayuseequivalentwording\.Forpreferencequestions,checkwhetheritsatisfiesthepersonalizationrubric\.Forknowledge\-updatequestions,accepttheupdatedanswerevenifolderinformationisalsopresent\.Fortemporal\-reasoningquestions,donotpenalizeanoff\-by\-oneerrorwhenthebenchmarkasksforanumberofdays,weeks,ormonths\.Forunanswerablequestions,answeryesonlywhentheresponsecorrectlyidentifiesthattherequestedinformationisnotgiven\.
Question:\{question\}
Referenceanswerorrubric:\{answer\}
Modelresponse:\{hypothesis\}
Answeryesornoonly\.
The LoCoMo implementation mapsCORRECTto one andWRONGto zero; the LongMemEval implementation mapsyesto one andnoto zero\. Both protocols use temperature zero\. The LoCoMo evaluator also fixes the seed to 42, while the LongMemEval checker uses top\-p=1p=1\. Category\-wise and overall accuracies are computed from these binary decisions\.相似文章
AdMem:面向任务求解智能体的高级记忆系统
本文介绍AdMem,一种面向基于LLM的智能体的统一记忆框架,整合语义记忆、情景记忆和程序性记忆,并采用双层短期与长期存储结构,通过多智能体架构实现自动记忆生成与自适应检索。实验表明,该方法在长程多轮任务中提升了鲁棒性和成功率。
DimMem:面向高效长期智能体记忆的维度结构化
DimMem 提出了一种用于 LLM 智能体的维度记忆框架,将记忆表示为具有显式字段的原子化、类型化单元,在 LoCoMo-10 和 LongMemEval-S 上实现了最先进的准确率,同时将 token 成本降低了 24%。
SelfMem: 面向AI智能体的自优化记忆框架
SelfMem提出了一种面向AI智能体的自优化记忆框架,使其能够通过记忆工具和反馈信号探索、评估和优化自身的记忆策略,在BEAM基准测试上,于大型对话规模下相较于基线方法取得了显著改进。
EverMemOS: 面向结构化长程推理的自组织记忆操作系统
EverMemOS 是一种面向大语言模型的自组织记忆操作系统,通过将对话结构化为记忆单元和场景来增强长程推理能力。
MESA:Task-Adaptive Multi-Structure Evidence Selection for Long-Horizon Agent Memory
This paper introduces MESA, a framework that dynamically selects and fuses a query-adaptive subset of multiple structural memory views for long-horizon agents, achieving 8.5% accuracy improvement over the strongest baseline while using 41% fewer evidence tokens.