记忆并非平等:面向 LLM 代理的分层协作记忆,实现有效性感知检索
摘要
文章提出了 HiCoMER,一个用于 LLM 代理的分层协作记忆管理框架,该框架通过提升有效性感知检索,减少过时记忆并增强问答质量。
查看缓存全文
缓存时间: 2026/09/28 09:33
# Not All Memories Are Equal: Hierarchical Collaborative Memory for Validity-Aware Retrieval in LLM Agents Source: [https://arxiv.org/html/2609.30289](https://arxiv.org/html/2609.30289) Rujing YaoAffiliation:Nanyang Technological University, SingaporeEmail:[rujing\.yao@ntu\.edu\.sg](mailto:[email protected])Ang LiAffiliation:University of Macau, ChinaEmail:[leeyonli@um\.edu\.mo](mailto:[email protected])Yang WuAffiliation:Worcester Polytechnic Institute, USAEmail:[ywu19@wpi\.edu](mailto:[email protected])Zhuoren Jiang\*Affiliation:Zhejiang University, ChinaEmail:[jiangzhuoren@zju\.edu\.cn](mailto:[email protected])Xiaozhong Liu\*Affiliation:Worcester Polytechnic Institute, USAEmail:[xliu14@wpi\.edu](mailto:[email protected]) ###### Abstract In team collaboration scenarios, memory is heterogeneous and continually evolving\. Team memories capture collective decisions, protocols, and current consensus, while individual memories preserve member\-specific observations, execution traces, and intermediate progress\. Existing memory\-augmented systems typically retrieve from all stored memories as a flat pool, ranking them by semantic relevance, importance, or recency without modeling hierarchical structure or evolving validity\. As a result, they often surface semantically relevant but outdated or conflicting memories, especially individual memories that no longer align with current team consensus, instead of prioritizing currently valid memories\. This is particularly problematic when collaborative LLM agents answer user questions, since their responses should be grounded in valid memories\. We propose HiCoMER, a framework for hierarchical collaborative memory management and validity\-aware retrieval in LLM agents\. HiCoMER first maintains the validity of team and individual memories and then retrieves memories that remain valid, rather than retrieving directly from all stored memories\. It consists of three components: a Hierarchical Memory Conflict Updater, a Validity\-Aware Memory Retriever, and a Memory\-Grounded Answer Generator\. To evaluate HiCoMER, we construct two new datasets for memory\-grounded question answering in collaborative settings\. Experiments on both datasets show that HiCoMER consistently outperforms strong baselines by reducing outdated retrieval, preserving current team consensus, and improving downstream QA quality\. 11footnotetext:Corresponding authors\.## 1Introduction Large language model \(LLM\) agents are increasingly moving beyond single\-turn assistance toward long\-horizon collaboration in shared environments[Boiko et al\. \(2023\)](https://arxiv.org/html/2609.30289#bib.bib1);[Gao et al\. \(2024\)](https://arxiv.org/html/2609.30289#bib.bib5)\. In these settings, memory is heterogeneous and hierarchical: team memories record collective decisions and protocols, while individual memories capture member\-specific observations and execution traces[Shuster et al\. \(2022\)](https://arxiv.org/html/2609.30289#bib.bib27)\. Figure 1:Illustration comparing existing methods and our proposed HiCoMER\.Agents may retrieve semantically relevant but invalid memories, overlooking current team consensus\. For example, when asked “How should we synthesize Compound A?”, an agent may select a past individual experiment log while ignoring a team\-level update banning Compound A due to toxicity\. Existing memory\-augmented systems[Lewis et al\. \(2020\)](https://arxiv.org/html/2609.30289#bib.bib2);[Packer et al\. \(2023\)](https://arxiv.org/html/2609.30289#bib.bib3)retrieve from a flat pool based on semantic similarity or recency, failing to capture hierarchical conflicts or evolving validity, as shown in Figure[1](https://arxiv.org/html/2609.30289#S1.F1)\. Moreover, collaborative memories are often produced in temporally aligned groups\. A single meeting or experiment cycle may generate both team\-level decisions and multiple individual records, some of which may conflict with the current consensus\. Memory management must therefore account for both within\-group hierarchical conflicts and cross\-time validity changes\. To address these challenges, we propose HiCoMER, a hierarchical collaborative memory framework\. HiCoMER maintains memory validity and performs validity\-aware retrieval through three components: a Hierarchical Memory Conflict Updater, a Validity\-Aware Memory Retriever, and a Memory\-Grounded Answer Generator\. This ensures agents prioritize currently valid memories\. We construct two new datasets for memory\-grounded question answering in collaborative settings\. Experiments show that HiCoMER consistently reduces outdated retrieval, preserves team consensus, and improves downstream QA quality\. Our contributions are as follows: - •We formalize the problem of hierarchical memory conflict in collaborative LLM settings, where both team and individual memories may evolve over time, and semantically relevant memories can become invalid under the current consensus\. - •We propose HiCoMER, a hierarchical collaborative memory framework that maintains memory validity and performs conflict\-aware, validity\-aware retrieval\. HiCoMER consists of three components: a Hierarchical Memory Conflict Updater, a Validity\-Aware Memory Retriever, and a Memory\-Grounded Answer Generator, which together enable agents to prioritize currently valid memories\. - •We construct two new datasets for memory\-grounded question answering in collaborative environments\. Experiments demonstrate that HiCoMER consistently improves retrieval safety, consensus preservation, and downstream QA performance over strong baselines\. ## 2Related Work ### 2\.1Long\-Term Memory Retrieval Retrieval\-augmented generation \(RAG\) grounds LLM outputs on retrieved evidence and is widely used in knowledge\-intensive NLP[Lewis et al\. \(2020\)](https://arxiv.org/html/2609.30289#bib.bib2)\. Many works improve the retrieval backbone, including dense retrievers for open\-domain QA such as DPR[Karpukhin et al\. \(2020\)](https://arxiv.org/html/2609.30289#bib.bib8), latent\-retrieval pretraining like REALM[Guu et al\. \(2020\)](https://arxiv.org/html/2609.30289#bib.bib9), and large\-scale retrieval\-augmented pretraining as in RETRO[Borgeaud et al\. \(2022\)](https://arxiv.org/html/2609.30289#bib.bib10)\. On the reader side, fusion\-based architectures such as FiD[Izacard and Grave \(2021\)](https://arxiv.org/html/2609.30289#bib.bib11)and end\-to\-end retrieval\-augmented models like Atlas[Izacard et al\. \(2023\)](https://arxiv.org/html/2609.30289#bib.bib12)enhance robustness and integration\. Complementary IR advances include late interaction models like ColBERT[Khattab and Zaharia \(2020\)](https://arxiv.org/html/2609.30289#bib.bib13), sparse expansion models like SPLADE[Formal et al\. \(2021\)](https://arxiv.org/html/2609.30289#bib.bib14), and query\-side augmentation such as HyDE[Gao et al\. \(2023b\)](https://arxiv.org/html/2609.30289#bib.bib15)\. Self\-reflective pipelines like Self\-RAG critique retrieved evidence and generations to improve faithfulness[Asai et al\. \(2024\)](https://arxiv.org/html/2609.30289#bib.bib28)\. Agent\-oriented systems treat memory as a persistent, growing corpus with write/read policies, e\.g\., MemGPT[Packer et al\. \(2023\)](https://arxiv.org/html/2609.30289#bib.bib3)and interactive generative agents[Park et al\. \(2023\)](https://arxiv.org/html/2609.30289#bib.bib4)\. Despite these advances, most approaches optimize flat relevance signals and do not explicitly model hierarchical validity or authority, which is central to our setting\. ### 2\.2Hierarchical Consistency under Conflicts Conflict resolution is often studied through contradiction detection, entailment reasoning, and factual verification\. NLI datasets like SNLI and MultiNLI provide supervision for entailment and contradiction[Bowman et al\. \(2015\)](https://arxiv.org/html/2609.30289#bib.bib16);[Williams et al\. \(2018\)](https://arxiv.org/html/2609.30289#bib.bib17), and adversarial benchmarks such as ANLI stress\-test robustness[Nie et al\. \(2020\)](https://arxiv.org/html/2609.30289#bib.bib18)\. Fact verification datasets like FEVER formalize consistency as verifying claims against evidence[Thorne et al\. \(2018\)](https://arxiv.org/html/2609.30289#bib.bib19)\. For generation, approaches include FactCC[Kryściński et al\. \(2020\)](https://arxiv.org/html/2609.30289#bib.bib20), QAGS\-style QA checks[Wang et al\. \(2020\)](https://arxiv.org/html/2609.30289#bib.bib21), and broader evaluations like TruthfulQA[Lin et al\. \(2022\)](https://arxiv.org/html/2609.30289#bib.bib22)and TRUE[Honovich et al\. \(2022\)](https://arxiv.org/html/2609.30289#bib.bib23)\. Post\-hoc detection and correction methods include SelfCheckGPT[Manakul et al\. \(2023\)](https://arxiv.org/html/2609.30289#bib.bib24), iterative retrieval\-revision pipelines like RARR[Gao et al\. \(2023a\)](https://arxiv.org/html/2609.30289#bib.bib25), and fine\-grained factual scoring such as FActScore[Min et al\. \(2023\)](https://arxiv.org/html/2609.30289#bib.bib26)\. However, these works mostly resolve conflicts at generation or evaluation time, rather than as write\-time operations on evolving memories\. In collaborative environments, conflicts are frequently asymmetric: team decisions and SOP updates can invalidate many individual logs even if semantically similar to future queries\. HiCoMER addresses this by organizing collaborative memories as time\-aligned groups, performing trainable maintenance over hierarchical and temporal conflicts, and learning a validity\-aware retrieval function over the maintained memory bank\. ## 3Methodology ### 3\.1Framework Overview Figure 2:Overall framework of HiCoMER\. The framework consists of three modules: Hierarchical Memory Conflict Updater, Validity\-Aware Memory Retriever, and Memory\-Grounded Answer Generator\.HiCoMER is a collaborative memory framework for long\-horizon LLM agents operating in team environments\. In collaborative settings, memory is inherently heterogeneous: Team memories capture shared decisions, protocols, and current coordination state, while Individual memories record local observations, execution traces, and intermediate progress\. Memory retrieval should therefore not rely solely on semantic relevance, because semantically relevant memories may already be outdated or conflict with other memories under the current collaborative state\. This issue becomes particularly severe when agents retrieve directly from raw stored memories without explicitly modeling the distinct functions of Team and Individual memories or maintaining the memory bank to resolve outdated and conflicting memories\. To address this problem, HiCoMER explicitly models collaborative memory at both the Team and Individual levels and maintains the memory bank to preserve a valid collaborative memory state for downstream question answering\. As illustrated in Figure[2](https://arxiv.org/html/2609.30289#S3.F2), HiCoMER consists of three modules\. The first module, Hierarchical Memory Conflict Updater, maintains the memory bank under two types of conflict: same\-time hierarchical conflict between Team and Individual memories in the current group, and cross\-time conflict between the current maintained Team/Individual memories and previously stored memories\. The second module, Validity\-Aware Memory Retriever, ranks candidate memories by combining semantic relevance with maintained memory\-state signals, so that retrieval is guided not only by textual relevance but also by whether a memory remains valid under the current collaborative state\. The third module, Memory\-Grounded Answer Generator, produces the final response based on the retrieved memories after a lightweight conflict\-resolution step\. Unlike conventional memory\-augmented systems that retrieve directly from raw stored memories, HiCoMER explicitly models Team and Individual memories and maintains their evolving validity over time\. This design allows the framework to reduce the retrieval of conflicting or outdated memories and ultimately generate answers grounded in currently valid memories rather than raw historical accumulation\. ### 3\.2Hierarchical Memory Conflict Updater This module maintains the collaborative memory bank by resolving outdated and conflicting memories under the hierarchical structure of Team and Individual memories\. Its goal is to preserve a valid collaborative memory state for downstream retrieval and answer generation\. At each time steptt, HiCoMER receives a time\-aligned memory group Gt=MtTeam∪MtIndividual,G\_\{t\}=M\_\{t\}^\{\\mathrm\{Team\}\}\\cup M\_\{t\}^\{\\mathrm\{Individual\}\},\(1\)whereMtTeamM\_\{t\}^\{\\mathrm\{Team\}\}contains collective records such as meeting resolutions, protocol updates, and supervisor instructions, whileMtIndividualM\_\{t\}^\{\\mathrm\{Individual\}\}contains individual records such as execution logs, observations, failed attempts, and intermediate progress\. During maintenance, each memory is associated with both its original content and its maintained content, so that HiCoMER preserves the historical record while updating the currently valid memory view used by downstream modules\. HiCoMER first identifies a bounded set of relevant historical memories, because the current group is unlikely to affect the entire historical memory bank and exhaustively comparing against all past memories would be computationally inefficient\. Specifically, it constructs a textual queryqthistq\_\{t\}^\{\\mathrm\{hist\}\}from the current Team and Individual memories and retrieves the Top\-KKmost relevant memories from the previously stored bank: Ht=TopKm∈ℳ<tsϕ\(qthist,m\),H\_\{t\}=\\operatorname\{TopK\}\_\{m\\in\\mathcal\{M\}\_\{<t\}\}\\;s\_\{\\phi\}\\\!\\left\(q\_\{t\}^\{\\mathrm\{hist\}\},m\\right\),\(2\)whereℳ<t\\mathcal\{M\}\_\{<t\}denotes all Team and Individual memories stored before timett,qthistq\_\{t\}^\{\\mathrm\{hist\}\}is constructed from the current groupGtG\_\{t\}, andsϕ\(⋅,⋅\)s\_\{\\phi\}\(\\cdot,\\cdot\)is a dense semantic retrieval model that computes the relevance between the current\-group query and each historical memory\.HtH\_\{t\}contains the Top\-KKhistorical Team/Individual memories with the highest retrieval scores\. Given the current groupGtG\_\{t\}and the retrieved historical candidatesHtH\_\{t\}, the updater performs maintenance at two levels\. The first level updates the current group itself\. WithinGtG\_\{t\}, Team and Individual memories may express inconsistent states\. HiCoMER therefore performs an intra\-group hierarchical update that revises the current Individual memories under the coordination signal provided by the current Team memories, yielding a maintained current\-group state: G~t=MtTeam∪M~tIndividual,\\widetilde\{G\}\_\{t\}=M\_\{t\}^\{\\mathrm\{Team\}\}\\cup\\widetilde\{M\}\_\{t\}^\{\\mathrm\{Individual\}\},\(3\)whereG~t\\widetilde\{G\}\_\{t\}denotes the maintained version of the current group after conflict resolution\. In this stage, the Team memories are treated as the authoritative coordination state at timett, and the updater revises Individual memories when they are incompatible with that state\. The second level updates the retrieved historical candidates\. HiCoMER uses the maintained current\-group stateG~t\\widetilde\{G\}\_\{t\}, which consists of the current Team memories and the updated current Individual memories, to perform an inter\-group hierarchical update over the historical candidates inHtH\_\{t\}\. This step revises previously stored Team and Individual memories whose maintained content has become outdated or conflicting under the current state\. We instantiate the updater with a lightweight instruction\-tuned LLM and train it to generate structured maintenance decisions rather than free\-form outputs\. For each target memory, the model predicts three elements: a relation label, an action label, and a revised maintained text when revision is required\. The relation label is drawn from\{compatible,neutral,contradiction\}\\\{\\mathrm\{compatible\},\\mathrm\{neutral\},\\mathrm\{contradiction\}\\\}, and the action label is drawn from\{keep,revise\}\\\{\\mathrm\{keep\},\\mathrm\{revise\}\\\}\. When the predicted action is “revise", HiCoMER updates the maintained content\. We train the updater in two stages\. We first perform supervised fine\-tuning on structured maintenance supervision, so that the model learns to produce the desired decision format and content\. We then further optimize it using Group Relative Policy Optimization \(GRPO\)[Shao et al\. \(2024\)](https://arxiv.org/html/2609.30289#bib.bib29)\. Given an input contextuu, the policy samples a set of candidate structured outputs 𝒴\(u\)=\{y~1,…,y~G\},\\mathcal\{Y\}\(u\)=\\\{\\tilde\{y\}\_\{1\},\\ldots,\\tilde\{y\}\_\{G\}\\\},\(4\)which are scored by a deterministic reward R\(y~,u\)=λ1Rcons\+λ2Rrev\+λ3Rfmt,R\(\\tilde\{y\};u\)=\\lambda\_\{1\}R\_\{\\mathrm\{cons\}\}\+\\lambda\_\{2\}R\_\{\\mathrm\{rev\}\}\+\\lambda\_\{3\}R\_\{\\mathrm\{fmt\}\},\(5\)whereRconsR\_\{\\mathrm\{cons\}\}measures consistency with the gold maintained state,RrevR\_\{\\mathrm\{rev\}\}encourages necessary but minimal revision, andRfmtR\_\{\\mathrm\{fmt\}\}enforces schema\-valid output\. In this way, GRPO encourages the updater to generate maintenance decisions that are accurate, concise, and structurally valid\. ### 3\.3Validity\-Aware Memory Retriever This module retrieves memories that are not only semantically relevant to the query but also valid under the current maintained collaborative state\. Its goal is to reduce the retrieval of outdated or conflicting memories while preserving evidence that remains useful for downstream answer generation\. Given a user queryqqand the maintained memory bankℳ≤t\\mathcal\{M\}\_\{\\leq t\}after processing groups up to timett, HiCoMER first retrieves a high\-recall candidate set Rt\(q\)=TopNm∈ℳ≤trψ\(q,m\),R\_\{t\}\(q\)=\\operatorname\{TopN\}\_\{m\\in\\mathcal\{M\}\_\{\\leq t\}\}\\;r\_\{\\psi\}\(q,m\),\(6\)whererψ\(⋅,⋅\)r\_\{\\psi\}\(\\cdot,\\cdot\)is a dense retrieval model that measures the semantic relevance between the query and each maintained memory, andRt\(q\)R\_\{t\}\(q\)contains the top\-NNcandidate memories returned by the first\-stage retriever\. HiCoMER then reranks each candidate memorym∈Rt\(q\)m\\in R\_\{t\}\(q\)using two complementary signals\. The first is a semantic relevance score Ssem\(q,m\)=fω\(q,m\),S\_\{\\mathrm\{sem\}\}\(q,m\)=f\_\{\\omega\}\(q,m\),\(7\)wherefω\(⋅,⋅\)f\_\{\\omega\}\(\\cdot,\\cdot\)is a cross\-encoder reranker that jointly encodes the query and the candidate memory for fine\-grained relevance scoring\. The second is a validity\-aware score that reflects whether a candidate memory remains usable under the current maintained state\. For each candidate memory, HiCoMER considers its maintained content, its source type \(Team or Individual\), and its most recent maintenance time\. Lethqh\_\{q\}andhmmainth\_\{m\}^\{\\mathrm\{maint\}\}denote the dense representations of the query and the maintained content of memorymm, respectively\. HiCoMER then constructs a feature vector v\(q,m\)=\[hmmaint;hq⊙hmmaint;e\(cm\);τm\],v\(q,m\)=\[h\_\{m\}^\{\\mathrm\{maint\}\};h\_\{q\}\\odot h\_\{m\}^\{\\mathrm\{maint\}\};e\(c\_\{m\}\);\\tau\_\{m\}\],\(8\)wheree\(cm\)e\(c\_\{m\}\)is a trainable embedding of the memory typecm∈\{Team,Individual\}c\_\{m\}\\in\\\{\\mathrm\{Team\},\\mathrm\{Individual\}\\\}, andτm\\tau\_\{m\}is a normalized timestamp feature derived from the latest maintenance time ofmm\. A lightweight MLP produces the validity\-aware score Sval\(q,m\)=gθ\(v\(q,m\)\)\.S\_\{\\mathrm\{val\}\}\(q,m\)=g\_\{\\theta\}\(v\(q,m\)\)\.\(9\) HiCoMER combines these two signals into a unified retrieval score: Sfinal\(q,m\)=Ssem\(q,m\)\+λ⋅logσ\(Sval\(q,m\)\),S\_\{\\mathrm\{final\}\}\(q,m\)=S\_\{\\mathrm\{sem\}\}\(q,m\)\+\\lambda\\cdot\\log\\sigma\\\!\\left\(S\_\{\\mathrm\{val\}\}\(q,m\)\\right\),\(10\)whereλ\\lambdacontrols the contribution of the validity\-aware component\. The final retrieved evidence is obtained by ranking candidate memories inRt\(q\)R\_\{t\}\(q\)according toSfinal\(q,m\)S\_\{\\mathrm\{final\}\}\(q,m\)\. We train the retriever after obtaining maintained memory states from Section[3\.2](https://arxiv.org/html/2609.30289#S3.SS2)\. For each queryqq, we construct pairwise ranking instances consisting of a positive memorym\+m^\{\+\}and a negative memorym−m^\{\-\}, wherem\+m^\{\+\}is a gold supporting memory that remains valid under the maintained state, whilem−m^\{\-\}is an outdated, conflicting, or semantically similar but non\-valid competitor\. The retriever is optimized with a pairwise ranking objective ℒretr=−log\(exp\(s\+\)exp\(s\+\)\+exp\(s−\)\)\.\\mathcal\{L\}\_\{\\mathrm\{retr\}\}=\-\\log\\left\(\\frac\{\\exp\(s^\{\+\}\)\}\{\\exp\(s^\{\+\}\)\+\\exp\(s^\{\-\}\)\}\\right\)\.\(11\)so that the validity\-aware scorer learns to rank currently valid memories above outdated or conflicting alternatives\. At inference time, HiCoMER first retrieves a candidate set using the dense retriever and then reranks the candidates using the unified scoreSfinal\(q,m\)S\_\{\\mathrm\{final\}\}\(q,m\)\. The resulting ranked memories are passed to the downstream answer generator\. ### 3\.4Memory\-Grounded Answer Generator This module generates the final answer from the retrieved memories\. Its goal is to produce responses grounded in the retrieved evidence while suppressing residual conflicts that may still remain among the top\-ranked candidates\. Given a user queryqqand the ranked memory set returned by the retriever, a lightweight conflict\-resolution step is applied over the retrieved evidence\. Although Module I and Module II already reduce outdated and conflicting memories, residual inconsistency may still remain when multiple memories are semantically relevant to the same query but reflect different collaborative states\. To address this issue, each retrieved memory is assigned a priority score based on its source type, maintained validity, and maintenance time: P\(m\)=α⋅𝕀\[cm=Team\]\+β⋅𝕀\[Valid\(m\)\]\+γ⋅NormTime\(tmedit\),\\begin\{split\}P\(m\)&=\\alpha\\cdot\\mathbb\{I\}\[c\_\{m\}=\\mathrm\{Team\}\]\\\\ &\\quad\+\\beta\\cdot\\mathbb\{I\}\[\\mathrm\{Valid\}\(m\)\]\\\\ &\\quad\+\\gamma\\cdot\\mathrm\{NormTime\}\(t\_\{m\}^\{edit\}\),\\end\{split\}\(12\)whereα\\alpha,β\\beta, andγ\\gammaare positive weights,cmc\_\{m\}denotes the memory type, andNormTime\(tmedit\)\\mathrm\{NormTime\}\(t\_\{m\}^\{edit\}\)is the normalized maintenance\-time score\. If two retrieved memories have contradictory maintained contents, the memory with higher priority is retained\. Detailed conflict\-resolution rules are provided in Appendix[A](https://arxiv.org/html/2609.30289#A1)\. The resulting evidence set is then passed to the answer generator\. The generator receives the query together with the resolved evidence set and produces the final response based solely on this evidence\. ## 4Experiments ### 4\.1Dataset Existing datasets lack explicit hierarchical conflicts and temporally\-aligned team versus individual memories, which are crucial for evaluating memory maintenance and validity\-aware retrieval in collaborative long\-horizon tasks\. To address this gap, we construct two new datasets corresponding to two biomedical collaboration settings: Acute Lung Injury \(ALI\) and PROTAC platform \(PROTAC\)\. The ALI dataset focuses on sepsis\-related therapeutic development, where conflicts are primarily driven by safety\-sensitive protocol updates, treatment discontinuation, and endpoint revisions\. The PROTAC dataset models linker\-free PROTAC platform development under N\-end rule constraints, with conflicts arising more often from target\-strategy revisions, assay replacements, and platform\-level design invalidations\. Table 1:Detailed statistics\.StatisticALIPROTACTotalCandidate papers303060Retained papers272653Valid groups9088731,781Memory documents8,9178,80117,718Injected conflicts317304621Same\-time conflicts189181370Cross\-time conflicts128123251Train papers191837Validation papers336Test papers5510 The detailed dataset statistics are summarized in Table[1](https://arxiv.org/html/2609.30289#S4.T1)\. The full construction process is provided in Appendix[B](https://arxiv.org/html/2609.30289#A2), and illustrative examples are provided in Appendix[C](https://arxiv.org/html/2609.30289#A3)\. ### 4\.2Baselines and Evaluation Metrics We compare HiCoMER against a diverse set of baselines covering standard semantic retrieval pipelines such as flat RAG, hybrid RAG, and rerank RAG; recency\-based memory control methods; agentic and hierarchical memory frameworks including MemGPT\-style agents[Packer et al\. \(2023\)](https://arxiv.org/html/2609.30289#bib.bib3)and G\-Memory[Zhang et al\. \(2025\)](https://arxiv.org/html/2609.30289#bib.bib7); and reflection\- or production\-oriented systems that rely on generation\-time reasoning without explicit memory\-bank maintenance, such as Self\-RAG[Asai et al\. \(2024\)](https://arxiv.org/html/2609.30289#bib.bib28)and Mem0[Chhikara et al\. \(2025\)](https://arxiv.org/html/2609.30289#bib.bib6)\. These baselines are chosen to test whether hierarchical conflict resolution can be addressed by stronger semantic matching alone, by recency heuristics, by existing long\-horizon memory mechanisms, or by generation\-time reflection without explicit memory maintenance\. We also include ablated variants of HiCoMER to isolate the contributions of the Hierarchical Memory Conflict Updater and the Validity\-Aware Memory Retriever\. For fairness, all methods are evaluated on the same grouped memory stream and query interface, and system\-level baselines are adapted to the same retrieval settings\. Full baseline definitions are provided in Appendix[D](https://arxiv.org/html/2609.30289#A4)\. Performance is evaluated at three levels\. For the Hierarchical Memory Conflict Updater, we measure structured maintenance quality using Conflict F1 and Maintenance Accuracy\. For upstream maintained\-state retrieval, we use Outdated Retrieval Rate at 5 \(ORR@5\) to quantify the reduction of outdated memory retrieval, Consensus Retention Rate at 5 \(CRR@5\) to evaluate preservation of authoritative Team consensus, and NDCG at 10 \(NDCG@10\) to measure overall executable ranking quality\. Finally, for the full HiCoMER pipeline, we report ROUGE\-L and Decision F1 to evaluate whether improvements in maintained\-state retrieval translate into more accurate and executable memory\-grounded answers\. Detailed metric definitions are provided in Appendix[D](https://arxiv.org/html/2609.30289#A4)\. ### 4\.3Implementation Details We evaluate HiCoMER on the ALI and PROTAC datasets under two backbone settings\. In the strong\-backbone setting, both the Hierarchical Memory Conflict Updater and Memory\-Grounded Answer Generator use Llama\-3\.1\-8B\. In the lightweight setting, both use Qwen2\.5\-3B\-Instruct\. The Validity\-Aware Memory Retriever is shared across settings\. All experiments use the same grouped memory streams, query interface, and downstream prompt templates to ensure fair comparison\. Training proceeds sequentially\. First, the updater is trained on structured maintenance instances to learn hierarchical memory conflict resolution\. The trained updater generates maintained memory states for the training split, which are then used to train the Validity\-Aware Memory Retriever with a pairwise ranking objective\. At inference, maintained memories are retrieved and reranked for executability, and the top memories are passed to the answer generator\. Table 2:End\-to\-end evaluation of the full HiCoMER pipeline on the ALI dataset using Llama\-3\.1\-8B as the backbone for the Hierarchical Memory Conflict Updater and the Memory\-Grounded Answer Generator\. Retrieval metrics ORR@5, CRR@5, and NDCG@10 measure maintained\-state retrieval quality, while answer\-level metrics ROUGE\-L and Decision F1 assess how well the pipeline converts maintained and retrieved memories into final memory\-grounded responses\. ### 4\.4Results We first present the main results on the ALI dataset, our primary in\-domain evaluation setting\. To assess effectiveness under a stronger backbone, we evaluate HiCoMER with Llama\-3\.1\-8B for both the Hierarchical Memory Conflict Updater and the Memory\-Grounded Answer Generator\. We also evaluate a lighter deployment\-oriented setting using Qwen2\.5\-3B\-Instruct to test the robustness of the framework under constrained compute and memory budgets\. Tables[2](https://arxiv.org/html/2609.30289#S4.T2)and[3](https://arxiv.org/html/2609.30289#S4.T3)report the end\-to\-end performance of HiCoMER\. Retrieval\-oriented metrics ORR@5, CRR@5, and NDCG@10 primarily measure maintained\-state retrieval quality, capturing the joint effect of the Hierarchical Memory Conflict Updater and the Validity\-Aware Memory Retriever\. Answer\-level metrics ROUGE\-L and Decision F1 further evaluate how improvements in retrieval translate into higher\-quality memory\-grounded responses\. Using Qwen2\.5\-3B\-Instruct, HiCoMER maintains strong performance, demonstrating that the maintenance\-and\-retrieval design is robust to smaller backbone models\. Gains are observed in both maintained\-state retrieval and final response quality\. Overall, HiCoMER substantially reduces outdated or non\-validity\-aware retrievals while better preserving authoritative Team consensus\. This leads to improved executable ranking and stronger downstream responses\. The consistent gains across both Llama\-3\.1\-8B and Qwen2\.5\-3B\-Instruct suggest that improvements stem from the framework design rather than backbone scale\. These results highlight that semantic relevance alone is insufficient in hierarchical collaborative memory settings: neither stronger semantic matching, recency heuristics, nor generation\-time reflection can fully replace explicit memory maintenance and validity\-aware retrieval\. Further component\-level evaluation of the Hierarchical Memory Conflict Updater on held\-out structured maintenance instances is provided in Appendix[E](https://arxiv.org/html/2609.30289#A5)\. Additional results on the PROTAC dataset and extended per\-baseline analyses are provided in Appendix[F](https://arxiv.org/html/2609.30289#A6)\. Table 3:End\-to\-end evaluation of HiCoMER on the ALI dataset using Qwen2\.5\-3B\-Instruct for both the Hierarchical Memory Conflict Updater and the Memory\-Grounded Answer Generator\. Retrieval metrics ORR@5, CRR@5, and NDCG@10 assess maintained\-state ranking, while answer\-level metrics ROUGE\-L and Decision F1 measure the quality of memory\-grounded responses under a smaller backbone\. ### 4\.5Ablation Study To quantify the contribution of each major module in HiCoMER, we conduct a full\-pipeline ablation study on the ALI dataset\. We remove each of the three core modules individually: the Hierarchical Memory Conflict Updater, the Validity\-Aware Memory Retriever, and the Memory\-Grounded Answer Generator\. Table[4](https://arxiv.org/html/2609.30289#S4.T4)reports both retrieval\-oriented metrics and answer\-level metrics to assess how degradations in upstream modules propagate to final response quality\. Table 4:Ablation study on the ALI dataset\. ORR@5, CRR@5, and NDCG@10 measure maintained\-state retrieval quality, while ROUGE\-L and Decision F1 assess memory\-grounded response quality\. Each row shows the effect of removing a single module from the full HiCoMER pipeline\.Removing either the updater or the validity\-aware retriever leads to substantial drops in both maintained\-state retrieval metrics and final response quality, confirming that these upstream modules are critical for end\-to\-end performance\. Removing the answer generator also degrades final response metrics, highlighting its role in converting retrieved memories into high\-quality answers\. Additional ablation analyses on the PROTAC dataset are provided in Appendix[G](https://arxiv.org/html/2609.30289#A7)\. ### 4\.6Human Evaluation We conduct a human evaluation to assess the practical performance of our proposed method\. Seventeen participants with research experience, including Ph\.D\. students, master’s students, and junior research assistants, evaluated 10 report\-writing tasks each, yielding a total of 170 evaluation instances\. For each task, participants evaluated outputs generated under the same scenario by HiCoMER and Mem0\. To reduce potential evaluation bias, the outputs were presented in randomized order, and participants were not informed which method produced each output\. We evaluate the outputs of HiCoMER and Mem0 in terms of two aspects, Memory Precision and Intervention Usefulness\. Memory Precision measures the relevance and accuracy of retrieved context, while Intervention Usefulness measures the actionability of the generated suggestions\. Participants rated each aspect on a 5\-point scale, with a score of 5 indicating the highest performance\. Table 5:Human evaluation results on simulated system utility, based on 170 evaluation instances\. Memory Precision measures the relevance and accuracy of retrieved context, and Intervention Usefulness measures the actionability of the generated suggestions\.The human evaluation results are reported in Table[5](https://arxiv.org/html/2609.30289#S4.T5)\. HiCoMER received higher ratings than Mem0 on both aspects of Memory Precision and Intervention Usefulness, suggesting that it provides more accurate memory support and more actionable suggestions in realistic scientific collaboration scenarios\. ## 5Conclusion We present HiCoMER, a hierarchical collaborative memory framework for long\-horizon LLM agents\. HiCoMER models team\-level and individual\-level information, maintains memory consistency with a trainable hierarchical conflict updater, retrieves currently valid memories via a validity\-aware scorer, and generates final answers grounded in the maintained evidence\. Experimental results on the datasets constructed in this work show that HiCoMER substantially reduces outdated retrieval while preserving valid memory states, improving downstream tasks and yielding higher\-quality, memory\-grounded answers and positive human evaluation outcomes\. Our results further suggest that effective long\-term memory management requires not only storing and retrieving historical information, but also identifying which memories remain valid as the collaborative context evolves\. Future work includes richer conflict structures, multi\-level supersession, and robustness under noisier or asynchronous memory streams\. ## Limitations While HiCoMER demonstrates consistent improvements in hierarchical memory maintenance and validity\-aware retrieval, a few limitations should be noted\. First, noisy, incomplete, or misaligned Team or Individual records could affect maintenance accuracy and downstream answer quality\. Finally, our current evaluation focuses on biomedical collaboration scenarios\. Future work will extend testing to additional domains\. ## Ethical Statement HiCoMER is designed for research and simulation of collaborative memory management in LLM agents and does not directly interact with human subjects or sensitive patient data in deployed settings\. All datasets used in this work are derived from publicly available scientific publications and are reverse\-simulated to avoid exposing identifiable human data\. For the human evaluation, participants were recruited voluntarily from research\-experienced individuals and performed scenario\-based tasks on simulated data only\. No private or sensitive information was used, and all participant data were anonymized\. Participants provided informed consent and were debriefed after the evaluation\. ## References - Asaiet al\.\(2024\)A\. Asai, Z\. Wu, Y\. Wang, A\. Sil, and H\. HajishirziSelf\-rag: learning to retrieve, generate, and critique through self\-reflection\.InInternational conference on learning representations,Vol\.2024,pp\. 9112–9141\.Cited by:[§2\.1](https://arxiv.org/html/2609.30289#S2.SS1.p1.1),[§4\.2](https://arxiv.org/html/2609.30289#S4.SS2.p1.1)\. - Boikoet al\.\(2023\)D\. A\. Boiko, R\. MacKnight, B\. Kline, and G\. GomesAutonomous chemical research with large language models\.Nature624\(7992\),pp\. 570–578\.Cited by:[§1](https://arxiv.org/html/2609.30289#S1.p1.1)\. - Borgeaudet al\.\(2022\)S\. Borgeaud, A\. Mensch, J\. Hoffmann, T\. Cai, E\. Rutherford, K\. Millican, G\. van den Driessche, J\. Lespiau, B\. Damoc, A\. Clark, D\. de Las Casas, A\. Guy, J\. Menick, R\. Ring, T\. Hennigan, S\. Huang, L\. Maggiore, C\. Jones, A\. Cassirer, A\. Brock, M\. Paganini, G\. Irving, O\. Vinyals, S\. Osindero, K\. Simonyan, J\. W\. Rae, E\. Elsen, and L\. SifreImproving language models by retrieving from trillions of tokens\.InInternational Conference on Machine Learning,pp\. 2206–2240\.Cited by:[§2\.1](https://arxiv.org/html/2609.30289#S2.SS1.p1.1)\. - Bowmanet al\.\(2015\)S\. R\. Bowman, G\. Angeli, C\. Potts, and C\. D\. ManningA large annotated corpus for learning natural language inference\.InProceedings of the 2015 conference on empirical methods in natural language processing,pp\. 632–642\.Cited by:[§2\.2](https://arxiv.org/html/2609.30289#S2.SS2.p1.1)\. - Chhikaraet al\.\(2025\)P\. Chhikara, D\. Khant, S\. Aryan, T\. Singh, and D\. YadavMem0: building production\-ready ai agents with scalable long\-term memory\.arXiv preprint arXiv:2504\.19413\.Cited by:[§4\.2](https://arxiv.org/html/2609.30289#S4.SS2.p1.1)\. - Formalet al\.\(2021\)T\. Formal, B\. Piwowarski, and S\. ClinchantSPLADE: sparse lexical and expansion model for first stage ranking\.InProceedings of the 44th international ACM SIGIR conference on research and development in information retrieval,pp\. 2288–2292\.Cited by:[§2\.1](https://arxiv.org/html/2609.30289#S2.SS1.p1.1)\. - Gaoet al\.\(2023a\)L\. Gao, Z\. Dai, P\. Pasupat, A\. Chen, A\. T\. Chaganty, Y\. Fan, V\. Zhao, N\. Lao, H\. Lee, D\. Juan, and K\. GuuRarr: researching and revising what language models say, using language models\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 16477–16508\.Cited by:[§2\.2](https://arxiv.org/html/2609.30289#S2.SS2.p1.1)\. - Gaoet al\.\(2023b\)L\. Gao, X\. Ma, J\. Lin, and J\. CallanPrecise zero\-shot dense retrieval without relevance labels\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 1762–1777\.Cited by:[§2\.1](https://arxiv.org/html/2609.30289#S2.SS1.p1.1)\. - Gaoet al\.\(2024\)S\. Gao, A\. Fang, Y\. Huang, V\. Giunchiglia, A\. Noori, J\. R\. Schwarz, Y\. Ektefaie, J\. Kondic, and M\. ZitnikEmpowering biomedical discovery with ai agents\.Cell187\(22\),pp\. 6125–6151\.Cited by:[§1](https://arxiv.org/html/2609.30289#S1.p1.1)\. - Guuet al\.\(2020\)K\. Guu, K\. Lee, Z\. Tung, P\. Pasupat, and M\. ChangRetrieval augmented language model pre\-training\.InInternational conference on machine learning,pp\. 3929–3938\.Cited by:[§2\.1](https://arxiv.org/html/2609.30289#S2.SS1.p1.1)\. - Honovichet al\.\(2022\)O\. Honovich, R\. Aharoni, J\. Herzig, H\. Taitelbaum, D\. Kukliansy, V\. Cohen, T\. Scialom, I\. Szpektor, A\. Hassidim, and Y\. MatiasTRUE: re\-evaluating factual consistency evaluation\.InProceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,pp\. 3905–3920\.Cited by:[§2\.2](https://arxiv.org/html/2609.30289#S2.SS2.p1.1)\. - Izacard and Grave \(2021\)G\. Izacard and E\. GraveLeveraging passage retrieval with generative models for open domain question answering\.InProceedings of the 16th conference of the european chapter of the association for computational linguistics: main volume,pp\. 874–880\.Cited by:[§2\.1](https://arxiv.org/html/2609.30289#S2.SS1.p1.1)\. - Izacardet al\.\(2023\)G\. Izacard, P\. Lewis, M\. Lomeli, L\. Hosseini, F\. Petroni, T\. Schick, J\. Dwivedi\-Yu, A\. Joulin, S\. Riedel, and E\. GraveAtlas: few\-shot learning with retrieval augmented language models\.Journal of Machine Learning Research24\(251\),pp\. 1–43\.Cited by:[§2\.1](https://arxiv.org/html/2609.30289#S2.SS1.p1.1)\. - Karpukhinet al\.\(2020\)V\. Karpukhin, B\. Oguz, S\. Min, P\. Lewis, L\. Wu, S\. Edunov, D\. Chen, and W\. YihDense passage retrieval for open\-domain question answering\.InProceedings of the 2020 conference on empirical methods in natural language processing \(EMNLP\),pp\. 6769–6781\.Cited by:[§2\.1](https://arxiv.org/html/2609.30289#S2.SS1.p1.1)\. - Khattab and Zaharia \(2020\)O\. Khattab and M\. ZahariaColbert: efficient and effective passage search via contextualized late interaction over bert\.InProceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval,pp\. 39–48\.Cited by:[§2\.1](https://arxiv.org/html/2609.30289#S2.SS1.p1.1)\. - Kryścińskiet al\.\(2020\)W\. Kryściński, B\. McCann, C\. Xiong, and R\. SocherEvaluating the factual consistency of abstractive text summarization\.InProceedings of the 2020 conference on empirical methods in natural language processing \(EMNLP\),pp\. 9332–9346\.Cited by:[§2\.2](https://arxiv.org/html/2609.30289#S2.SS2.p1.1)\. - Lewiset al\.\(2020\)P\. Lewis, E\. Perez, A\. Piktus, F\. Petroni, V\. Karpukhin, N\. Goyal, H\. Küttler, M\. Lewis, W\. Yih, T\. Rocktäschel, S\. Riedel, and D\. KielaRetrieval\-augmented generation for knowledge\-intensive NLP tasks\.InAdvances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020,Cited by:[§1](https://arxiv.org/html/2609.30289#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.30289#S2.SS1.p1.1)\. - Linet al\.\(2022\)S\. Lin, J\. Hilton, and O\. EvansTruthfulqa: measuring how models mimic human falsehoods\.InProceedings of the 60th annual meeting of the association for computational linguistics \(volume 1: long papers\),pp\. 3214–3252\.Cited by:[§2\.2](https://arxiv.org/html/2609.30289#S2.SS2.p1.1)\. - Manakulet al\.\(2023\)P\. Manakul, A\. Liusie, and M\. GalesSelfcheckgpt: zero\-resource black\-box hallucination detection for generative large language models\.InProceedings of the 2023 conference on empirical methods in natural language processing,pp\. 9004–9017\.Cited by:[§2\.2](https://arxiv.org/html/2609.30289#S2.SS2.p1.1)\. - Minet al\.\(2023\)S\. Min, K\. Krishna, X\. Lyu, M\. Lewis, W\. Yih, P\. Koh, M\. Iyyer, L\. Zettlemoyer, and H\. HajishirziFActScore: fine\-grained atomic evaluation of factual precision in long form text generation\.InProceedings of the 2023 conference on empirical methods in natural language processing,pp\. 12076–12100\.Cited by:[§2\.2](https://arxiv.org/html/2609.30289#S2.SS2.p1.1)\. - Nieet al\.\(2020\)Y\. Nie, A\. Williams, E\. Dinan, M\. Bansal, J\. Weston, and D\. KielaAdversarial nli: a new benchmark for natural language understanding\.InProceedings of the 58th annual meeting of the association for computational linguistics,pp\. 4885–4901\.Cited by:[§2\.2](https://arxiv.org/html/2609.30289#S2.SS2.p1.1)\. - Packeret al\.\(2023\)C\. Packer, S\. Wooders, K\. Lin, V\. Fang, S\. G\. Patil, I\. Stoica, and J\. E\. GonzalezMemgpt: towards llms as operating systems\.arXiv preprint arXiv:2310\.08560\.Cited by:[§1](https://arxiv.org/html/2609.30289#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.30289#S2.SS1.p1.1),[§4\.2](https://arxiv.org/html/2609.30289#S4.SS2.p1.1)\. - Parket al\.\(2023\)J\. S\. Park, J\. C\. O’Brien, C\. J\. Cai, M\. R\. Morris, P\. Liang, and M\. S\. BernsteinGenerative agents: interactive simulacra of human behavior\.InProceedings of the 36th Annual ACM Symposium on User Interface Software and Technology,pp\. 1–22\.Cited by:[§2\.1](https://arxiv.org/html/2609.30289#S2.SS1.p1.1)\. - Shaoet al\.\(2024\)Z\. Shao, P\. Wang, Q\. Zhu, R\. Xu, J\. Song, X\. Bi, H\. Zhang, M\. Zhang, Y\. K\. Li, Y\. Wu, and D\. GuoDeepSeekMath: pushing the limits of mathematical reasoning in open language models\.arXiv preprint arXiv:2402\.03300\.Cited by:[§3\.2](https://arxiv.org/html/2609.30289#S3.SS2.p7.1)\. - Shusteret al\.\(2022\)K\. Shuster, J\. Xu, M\. Komeili, D\. Ju, E\. M\. Smith, S\. Roller, M\. Ung, M\. Chen, K\. Arora, J\. Lane, M\. Behrooz, W\. Ngan, S\. Poff, N\. Goyal, A\. Szlam, Y\. Boureau, M\. Kambadur, and J\. WestonBlenderbot 3: a deployed conversational agent that continually learns to responsibly engage\.arXiv preprint arXiv:2208\.03188\.Cited by:[§1](https://arxiv.org/html/2609.30289#S1.p1.1)\. - Thorneet al\.\(2018\)J\. Thorne, A\. Vlachos, C\. Christodoulopoulos, and A\. MittalFEVER: a large\-scale dataset for fact extraction and verification\.InProceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long Papers\),pp\. 809–819\.Cited by:[§2\.2](https://arxiv.org/html/2609.30289#S2.SS2.p1.1)\. - Wanget al\.\(2020\)A\. Wang, K\. Cho, and M\. LewisAsking and answering questions to evaluate the factual consistency of summaries\.InProceedings of the 58th annual meeting of the association for computational linguistics,pp\. 5008–5020\.Cited by:[§2\.2](https://arxiv.org/html/2609.30289#S2.SS2.p1.1)\. - Williamset al\.\(2018\)A\. Williams, N\. Nangia, and S\. R\. BowmanA broad\-coverage challenge corpus for sentence understanding through inference\.InProceedings of the 2018 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 \(long papers\),pp\. 1112–1122\.Cited by:[§2\.2](https://arxiv.org/html/2609.30289#S2.SS2.p1.1)\. - Zhanget al\.\(2025\)G\. Zhang, M\. Fu, K\. Wang, F\. Wan, M\. Yu, and S\. YanG\-memory: tracing hierarchical memory for multi\-agent systems\.InAdvances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025,Cited by:[§4\.2](https://arxiv.org/html/2609.30289#S4.SS2.p1.1)\. ## Appendix ADetailed Conflict Resolution in Answer Generation Given two retrieved memoriesmim\_\{i\}andmjm\_\{j\}, HiCoMER checks if their maintained contents are contradictory: Contradict\(xmimaint,xmjmaint\)=1\.\\mathrm\{Contradict\}\(x\_\{m\_\{i\}\}^\{\\mathrm\{maint\}\},x\_\{m\_\{j\}\}^\{\\mathrm\{maint\}\}\)=1\.\(13\) If no contradiction is detected, both memories are retained\. If a contradiction exists, the memory with the higher priority is retained: mi≻mjifContradict\(xmimaint,xmjmaint\)=1∧P\(mi\)\>P\(mj\)\.\\begin\{split\}m\_\{i\}\\succ m\_\{j\}\\quad\\text\{if\}\\quad&\\mathrm\{Contradict\}\(x\_\{m\_\{i\}\}^\{\\mathrm\{maint\}\},x\_\{m\_\{j\}\}^\{\\mathrm\{maint\}\}\)=1\\\\ &\\land\\ P\(m\_\{i\}\)\>P\(m\_\{j\}\)\.\\end\{split\}\(14\) The final evidence set for answer generation is then defined as 𝒞∗=\{m∈Rt\(q\)∣∄m′∈Rt\(q\)such thatm′≻m\}\.\\begin\{split\}\\mathcal\{C\}^\{\*\}=\\\{m\\in R\_\{t\}\(q\)\\ \\mid\\ &\\nexists\\,m^\{\\prime\}\\in R\_\{t\}\(q\)\\ \\text\{such that\}\\\\ &m^\{\\prime\}\\succ m\\\}\.\\end\{split\}\(15\) All else being equal, this procedure yields the following preference order in common cases: - •Team memories are preferred over Individual memories when their maintained contents conflict\. - •Among memories of the same type, those that remain valid under the maintained state are preferred over invalid or outdated ones\. - •If both memories have the same type and validity, the memory with the more recent maintenance time is preferred\. After conflict resolution, each memory in𝒞∗\\mathcal\{C\}^\{\*\}is serialized with its source type, maintenance time, and maintained content, and the resulting structured evidence set is passed to the instruction\-following LLM for final answer generation\. ## Appendix BALI and PROTAC Dataset Construction Details This appendix provides full details on the construction of the ALI and PROTAC datasets\. For each domain, we start with a candidate pool of 30 PubMed papers\. After paper\-level eligibility screening, 27 ALI papers and 26 PROTAC papers are retained\. Each retained paper is reverse\-simulated into a 12\-week collaborative lifecycle, divided into event windows spanning 48–72 hours\. Each event window corresponds to a grouped memory unit containing Team memories, which encode collective decisions, protocol\-level constraints, and authoritative consensus, and Individual memories, which capture local execution traces, observations, and intermediate progress\. Hierarchical conflicts are explicitly injected, including same\-time conflicts within a single event window and cross\-time conflicts spanning multiple windows\. ALI contains 317 injected conflicts \(189 same\-time, 128 cross\-time\), and PROTAC contains 304 injected conflicts \(181 same\-time, 123 cross\-time\)\. To reduce document\-level leakage, train/validation/test splits are organized by source paper: ALI uses 19/3/5 and PROTAC uses 18/3/5\. Both datasets support six query types: fact lookup, procedure validation, conflict diagnosis, risk warning, next\-step decision, and protocol consistency checking\. GPT\-5 is used for reverse simulation, query\-answer generation, and evidence\-trace annotation\. Across valid groups, each contains on average 2\.1 Team memories and 7\.8 Individual memories, totaling 9\.9 documents per group\. ## Appendix CIllustrative examples Illustrative examples for the ALI and PROTAC datasets are provided in Table[6](https://arxiv.org/html/2609.30289#A3.T6)\. Table 6:Illustrative examples\.Table 7:Component\-level evaluation of the Hierarchical Memory Conflict Updater on held\-out structured maintenance instances from ALI\. Conflict F1 measures the accuracy of predicting conflict relations, while Maintenance Accuracy measures the correctness of structured memory updates\. ## Appendix DBaseline Definitions and Metric Details To evaluate HiCoMER’s handling of hierarchical memory conflicts, we compare it against several baseline categories\. All baselines are evaluated on the same grouped memory stream and query interface to ensure differences reflect memory maintenance and retrieval strategies rather than input format or generation capacity\. - •Standard retrieval and semantic matching:Flat RAG ranks candidates by dense semantic similarity\. Hybrid RAG combines dense and sparse \(BM25\) retrieval\. Rerank RAG applies a cross\-encoder to rerank Hybrid RAG candidates\. These provide strong semantic matching baselines\. - •Recency\-based memory control:MemGPT\-style Window retains only the most recent memories within a fixed active window, simulating recency\-focused memory management\. - •Agentic and hierarchical memory systems:Agentic Memory scores candidates by relevance, recency, and static importance\. G\-Memory organizes memories hierarchically and performs structured retrieval over the shared memory pool\. - •Reflection and production memory frameworks:Self\-RAG applies generation\-time reflection to assess whether retrieved memories remain valid\. Mem0 is a production memory infrastructure adapted to the shared\-memory benchmark\. - •Internal HiCoMER variants:HiCoMER without validity\-aware retrieval retains the maintenance module but ranks memories solely by semantic similarity\. - •Updater\-only baselines:Rule\-based Updater applies deterministic revise rules\. Qwen2\.5\-3B Zero\-shot uses direct structured prompting without training\. Qwen2\.5\-3B \+ SFT applies supervised fine\-tuning\. Qwen2\.5\-3B \+ SFT \+ GRPO is the full trainable updater\. Llama\-3\.1\-8B Zero\-shot/SFT/SFT\+GRPO provide larger\-backbone variants for robustness evaluation\. Metrics are defined to assess multiple dimensions of performance: - •Updater\-specific metrics:Conflict F1 measures accuracy in predicting conflict relations\. Maintenance Accuracy measures exact\-match correctness of relation and action labels\. - •Retrieval\-oriented metrics:ORR@K \(Outdated Retrieval Rate\) is the fraction of retrieved memories labeled outdated or non\-executable\. CRR@K \(Consensus Retention Rate\) is the recall of authoritative Team memories\. NDCG@10 evaluates ranking quality with higher gain assigned to valid supporting memories\. - •Answer\-level full\-pipeline metrics:ROUGE\-L measures textual overlap between generated and gold answers\. Decision F1 evaluates correctness on decision\-oriented queries, including risk warnings, next\-step decisions, and protocol consistency checks\. Conflict F1 and Maintenance Accuracy primarily assess the updater, ORR@5, CRR@5, and NDCG@10 reflect the combined effect of the updater and the validity\-aware memory retriever, and ROUGE\-L and Decision F1 evaluate the end\-to\-end quality of memory\-grounded answer generation\. ## Appendix EUpdater Evaluation We evaluate the Hierarchical Memory Conflict Updater on held\-out structured maintenance instances derived from the ALI dataset\. This component\-level evaluation measures whether the updater can correctly predict conflict relations and structured maintenance actions\. The results show that the trained updater substantially outperforms rule\-based maintenance, indicating that hierarchical memory updating cannot be reduced to standard contradiction detection alone\. Reward\-based GRPO optimization further improves structured executable\-state maintenance, yielding consistent gains over zero\-shot and SFT\-only settings\. Table 8:End\-to\-end evaluation of the full HiCoMER pipeline on the PROTAC dataset using Llama\-3\.1\-8B as the backbone for the Hierarchical Memory Conflict Updater and the Memory\-Grounded Answer Generator\. Retrieval metrics ORR@5, CRR@5, and NDCG@10 measure maintained\-state retrieval quality, while answer\-level metrics ROUGE\-L and Decision F1 assess how well the pipeline converts maintained and retrieved memories into final memory\-grounded responses\. ## Appendix FResults on PROTAC Dataset Table 9:End\-to\-end evaluation of HiCoMER on the PROTAC dataset using Qwen2\.5\-3B\-Instruct for both the Hierarchical Memory Conflict Updater and the Memory\-Grounded Answer Generator\. Retrieval metrics ORR@5, CRR@5, and NDCG@10 assess maintained\-state ranking, while answer\-level metrics ROUGE\-L and Decision F1 measure the quality of memory\-grounded responses under a smaller backbone\.We present the full main results on the PROTAC dataset in this appendix\. As in the main text, the tables report end\-to\-end evaluations of the complete HiCoMER pipeline\. Metrics ORR@5, CRR@5, and NDCG@10 primarily capture the combined effect of the Hierarchical Memory Conflict Updater and Validity\-Aware Memory Retriever, while ROUGE\-L and Decision F1 assess the final response quality produced by the Memory\-Grounded Answer Generator\. Table 10:Ablation study on the PROTAC dataset\. ORR@5, CRR@5, and NDCG@10 measure maintained\-state retrieval quality, while ROUGE\-L and Decision F1 assess memory\-grounded response quality\. Each row shows the effect of removing a single module from the full HiCoMER pipeline\.The evaluation on the PROTAC dataset demonstrates that HiCoMER consistently reduces outdated retrieval and preserves valid memory states under complex, evolving workflow constraints\. Improvements in maintained\-state retrieval translate into higher\-quality, memory\-grounded responses, reflecting the effective coordination of hierarchical memory maintenance and validity\-aware retrieval\. Overall, these results indicate that HiCoMER provides a robust and effective framework for managing dynamic collaborative memories in mechanism\-driven platform settings\. ## Appendix GAblation Study on the PROTAC Dataset To further assess module contributions across datasets, we conduct an ablation study on the PROTAC dataset\. Each of the three core HiCoMER modules is removed individually: the Hierarchical Memory Conflict Updater, the Validity\-Aware Memory Retriever, and the Memory\-Grounded Answer Generator\. Table[10](https://arxiv.org/html/2609.30289#A6.T10)reports retrieval\-oriented and answer\-level metrics to evaluate how module removals affect maintained\-state retrieval and final response quality\. Removing either the updater or the validity\-aware retriever results in significant drops across both retrieval and answer\-level metrics, confirming that these upstream modules are essential for end\-to\-end performance\. Removing the answer generator also reduces response quality, demonstrating its role in converting retrieved memories into high\-quality answers\. These results complement the ALI dataset ablation study and show consistent module contributions across datasets\.
相似文章
HeLa-Mem:面向LLM智能体的赫布学习与联想记忆
# HeLa-Mem: Hebbian Learning and Associative Memory for LLM Agents 来源:[https://arxiv.org/html/2604.16839](https://arxiv.org/html/2604.16839) Jinchang Zhu1,∗,a, Jindong Li1,∗, Cheng Zhang2,∗, Jiahong Liu3, Menglin Yang1,†,b 1香港科技大学(广州) 2吉林大学 3香港中文大学 [email protected] [email protected] ∗同等贡献 †通讯作者 ###### 摘要 长...
H-Mem:一种通过混合结构实现智能体记忆演化与检索的新型记忆机制
H-Mem是一种面向基于LLM的智能体的新型记忆机制,采用时间-语义树与知识图谱相结合的混合结构,以建模记忆演化并提升检索性能,在问答基准上实现了最先进水平。
面向LLM智能体的层级图记忆:路径级定位与重写
本文介绍了HiGram,一种面向LLM智能体的演化式层级图记忆框架,具备路径级定位与协同重写功能,以提高长期推理任务中的检索效率和答案质量。
面向长期LLM代理的检索驱动记忆再巩固
本文介绍了REALM,一个用于LLM代理长期记忆的框架,它利用检索驱动的再巩固来自主地将记忆组织成认知图,从而在长期记忆基准测试中取得更好的性能。
STALE:LLM智能体能否识别记忆何时失效?
本文识别了LLM智能体中的一个关键失效模式:当新证据与先前信念冲突时,它们无法更新个性化记忆。本文引入了STALE基准和一个三维探测框架,揭示了即使最佳模型也仅达到55.2%的准确率,并提出了CUPMem作为鲁棒记忆修正的原型。