To Memories and Beyond: From Remembering to Knowing You across Long-Term Multimodal Personal Archives

arXiv cs.CL Papers

Summary

This paper introduces ReaLMem, the first benchmark for evaluating long-term multimodal memory in AI using authentic personal archives, and proposes ChronoProfiler for temporal weighting to improve personalization in AI companions.

arXiv:2609.19167v1 Announce Type: new Abstract: As AI systems evolve into personalized digital companions, a central capability is reasoning over a user's long-term personal history: not merely storing past events, but tracking longitudinal experiences and evolving preferences. Progress here is bottlenecked by evaluation, existing long-term memory benchmarks are largely synthetic and text-only, they overlook the visual records that anchor everyday human memory, lack the authentic and causally connected longitudinal data that real personalization demands, and consequently remain confined to shallow factual recall. We introduce ReaLMem (Real-world Long-term Multimodal Memory), the first benchmark built from authentic multi-year personal visual archives, paired with first-person subjective annotations. ReaLMem evaluates models across three cognitive tiers of increasing difficulty: factual recall, persona inference, and predictive personalization. We further propose ChronoProfiler, a temporal-weighting profiling module that computes temporal stability scores for user attributes and applies them as a salience prior, resolving conflicts among temporally inconsistent preferences and helping models compound multiple co-active preferences in complex personalized decisions. Extensive evaluation of frontier multimodal large language models (MLLMs) and memory systems on ReaLMem reveals predictive personalization as a consistent ceiling, exposes clear performance gaps and bottlenecks between MLLMs and memory systems, and shows that high-quality, temporally informed representations substantially improve personalization. Together, ReaLMem and ChronoProfiler provide an authentic testbed and a simple, effective mechanism for long-term personalization, laying a foundation for future research on lifelong AI companions.
Original Article
View Cached Full Text

Cached at: 09/18/26, 08:53 AM

# To Memories and Beyond: From Remembering to Knowing You across Long-Term Multimodal Personal Archives
Source: [https://arxiv.org/html/2609.19167](https://arxiv.org/html/2609.19167)
Zhuorui YuAffiliation:Memories\.ai ResearchKaiao WenAffiliation:University of BristolHao ZhengAffiliation:University of BristolAffiliation:Memories\.ai ResearchXinyi ZhengAffiliation:University of BristolPeiran WuAffiliation:University of BristolAffiliation:Memories\.ai ResearchEnmin ZhouAffiliation:Memories\.ai ResearchChi\-Hao WuJunxiao ShenAffiliation:University of Bristol

###### Abstract

As AI systems evolve into personalized digital companions, a central capability is reasoning over a user’s long\-term personal history: not merely storing past events, but tracking longitudinal experiences and evolving preferences\. Progress here is bottlenecked by evaluation, existing long\-term memory benchmarks are largely synthetic and text\-only, they overlook the visual records that anchor everyday human memory, lack the authentic and causally connected longitudinal data that real personalization demands, and consequently remain confined to shallow factual recall\. We introduce ReaLMem \(Real\-world Long\-term Multimodal Memory\), the first benchmark built from authentic multi\-year personal visual archives, paired with first\-person subjective annotations\. ReaLMem evaluates models across three cognitive tiers of increasing difficulty: factual recall, persona inference, and predictive personalization\. We further propose ChronoProfiler, a temporal\-weighting profiling module that computes temporal stability scores for user attributes and applies them as a salience prior, resolving conflicts among temporally inconsistent preferences and helping models compound multiple co\-active preferences in complex personalized decisions\. Extensive evaluation of frontier multimodal large language models \(MLLMs\) and memory systems on ReaLMem reveals predictive personalization as a consistent ceiling, exposes clear performance gaps and bottlenecks between MLLMs and memory systems, and shows that high\-quality, temporally informed representations substantially improve personalization\. Together, ReaLMem and ChronoProfiler provide an authentic testbed and a simple, effective mechanism for long\-term personalization, laying a foundation for future research on lifelong AI companions\.

22footnotetext:Corresponding author\.## 1Introduction

As AI evolves toward Artificial General Intelligence \(AGI\), human–computer interaction is shifting from one\-off instruction execution toward persistent, personalized digital companions\. Long\-term memory is the cornerstone of this shift\[[22](https://arxiv.org/html/2609.19167#bib.bib1),[46](https://arxiv.org/html/2609.19167#bib.bib2)\]: it lets a model accumulate, track, and reason over a user’s experiences and evolving preferences across years, moving beyond isolated factual Q&A toward responses tailored to an individual’s history\. Accordingly, recent work has begun to build benchmarks for long\-context, long\-horizon personal memory\[[43](https://arxiv.org/html/2609.19167#bib.bib3),[29](https://arxiv.org/html/2609.19167#bib.bib4),[39](https://arxiv.org/html/2609.19167#bib.bib29)\]\. Yet progress toward genuinely personalized long\-term assistants is held back by two limitations of current evaluation—the data and the tasks\.

![Refer to caption](https://arxiv.org/html/2609.19167v1/task_taxonomy.png)Figure 1:Overview of ReaLMem: \(left\) example personal multimodal memory stream over years; \(right\) the three\-tier task pyramid grounded in this memory, rising in cognitive complexity from Factual Memory Recall \(Tier 1\) and Persona Inference \(Tier 2\) to Predictive Personalization \(Tier 3\), with each of the nine subtasks shown alongside its definition and an example query\.First, current benchmarks are largely inauthentic and vision\-blind\. Constrained by privacy and cost, they are predominantly built top\-down from text\-only data synthesized by large language models \(LLMs\)\[[48](https://arxiv.org/html/2609.19167#bib.bib6),[7](https://arxiv.org/html/2609.19167#bib.bib5),[18](https://arxiv.org/html/2609.19167#bib.bib7),[19](https://arxiv.org/html/2609.19167#bib.bib8)\], or stretched to long context by mechanically inserting irrelevant passages \(*e\.g*\., “needle\-in\-a\-haystack” padding\)\[[14](https://arxiv.org/html/2609.19167#bib.bib18),[36](https://arxiv.org/html/2609.19167#bib.bib26)\]\. Such synthesis disrupts the natural temporal and causal structure of memory and rarely captures the messiness of genuine human experience\. More fundamentally, it overlooks the visual record—the photos and videos that anchor much of everyday human memory\[[2](https://arxiv.org/html/2609.19167#bib.bib31)\]—leaving evaluation disconnected from the physical world\.

Second, as a direct consequence, evaluation stops at shallow factual recall\. Because synthetic corpora lack the authentic, causally connected longitudinal data that accrues in real life, they cannot support the deeper competencies that genuine personalization demands\. Human preferences are dynamic and context\-dependent, shaped by the interplay of situation and emotion; assessing personalized assistants therefore requires moving beyond retrieval toward understanding how preferences evolve, reasoning about why they change, and predicting future choices\.

These limitations raise two questions: \(a\) what data can ground the evaluation of authentic long\-term personalization; \(b\) how should tasks be designed to comprehensively evaluate the ability to understand human personalization?

Egocentric and lifelog vision\[[10](https://arxiv.org/html/2609.19167#bib.bib32),[6](https://arxiv.org/html/2609.19167#bib.bib49),[4](https://arxiv.org/html/2609.19167#bib.bib33)\]has begun to study personal visual data directly, but mainly for short\-horizon perception and retrieval, and typically without the data owner’s first\-person subjective annotations—two gaps that have kept personalization underdeveloped\. We argue that the personal media archives already stored on people’s devices—especially photos and videos—are instead a natural and powerful substrate for*long\-term*multimodal memory: spanning years, they record where users went, what they did, and what mattered to them, preserving the visual and contextual anchors that text\-only logs discard\. Treated as long\-term memory and paired with the owner’s own perspective, these archives open a path to assistants that provide proactive, context\-aware, and personalized help grounded in visual memory\.

Building on above insights, we depart from the synthetic\-data paradigm and introduce ReaLMem, the first multimodal long\-term memory benchmark dataset derived from authentic personal visual archives paired with first\-person subjective annotations\. To structure evaluation, and inspired by Bloom’s taxonomy of educational objectives\[[1](https://arxiv.org/html/2609.19167#bib.bib24)\], we organize tasks into three progressively harder cognitive tiers \(Fig\.[1](https://arxiv.org/html/2609.19167#S1.F1)\): from factual recall, through persona inference, to predictive personalization over a user’s history\.

Realistic personalization at this scale also faces a concrete architectural obstacle: feeding ultra\-long multimodal histories to a model is computationally costly and prone to the lost\-in\-the\-middle effect\[[28](https://arxiv.org/html/2609.19167#bib.bib17)\], while existing approaches flatten historical preferences\[[38](https://arxiv.org/html/2609.19167#bib.bib20),[42](https://arxiv.org/html/2609.19167#bib.bib30)\], treating one\-off mentions and long\-held habits alike and leaving no principled way to arbitrate when preferences conflict or compete\. As a targeted remedy, we propose ChronoProfiler, a plug\-and\-play user\-profiling module that incrementally distills a structured persona from longitudinal history and assigns each item a temporal stability score\. Used as a salience prior, this score resolves temporal conflicts and helps the model weigh and compound co\-active preferences in complex personalized decisions\.

In summary, our contributions are:

∙\\bulletWe introduce ReaLMem, the first multimodal long\-term memory benchmark built from authentic personal visual archives and first\-person subjective annotations, comprising 2,508 multimodal sessions spanning more than six years with genuine spatio\-temporal metadata\.

∙\\bulletWe design a three\-tier cognitive evaluation hierarchy—Factual Memory Recall, Persona Inference, and Predictive Personalization—instantiated by 1,629 QA pairs for fine\-grained assessment of personalization\.

∙\\bulletWe propose ChronoProfiler, a temporal\-aware user\-profiling module that turns a temporal stability score into a salience prior for preference\-weighted personalized tasks\.

∙\\bulletWe comprehensively evaluate multimodal LLMs \(MLLMs\) and memory systems on ReaLMem, revealing performance gaps, modality dependencies, and open challenges\.

## 2Related Works

### 2\.1Personal Memory Benchmarks

Recent personal memory benchmarks\[[48](https://arxiv.org/html/2609.19167#bib.bib6),[45](https://arxiv.org/html/2609.19167#bib.bib21),[17](https://arxiv.org/html/2609.19167#bib.bib27),[24](https://arxiv.org/html/2609.19167#bib.bib28)\]approach the long\-term memory challenge along three complementary trajectories\. A first line broadens modality and source coverage by generating multi\-session, multimodal dialogues from causal temporal event graphs\[[29](https://arxiv.org/html/2609.19167#bib.bib4)\]\. A second line embeds user preferences implicitly within conversational context, testing deductive reasoning over fragmented and incidental signals\[[7](https://arxiv.org/html/2609.19167#bib.bib5),[18](https://arxiv.org/html/2609.19167#bib.bib7),[19](https://arxiv.org/html/2609.19167#bib.bib8)\]\. A third line scales context length via a “needle\-in\-a\-haystack” construction that injects critical evidence into vast amounts of distractor dialogues\[[43](https://arxiv.org/html/2609.19167#bib.bib3)\], reaching contexts of up to 1\.5 M tokens\.

Despite their progress, these benchmarks share three structural limitations rooted in their synthetic origin\. \(i\) Modality narrowness, or vision\-blindness: most evaluation data is text\-only, leaving the visual signals that anchor everyday human memory under\-explored\. \(ii\) Density loss and structural flattening: persona templates and temporal graphs, when realized through LLM synthesis, erase the modality cues, causal continuity, and natural temporal alignment of lived experience, producing dialogues that lack genuine contextual nuance and unpredictability\. \(iii\) Limited cognitive depth: as a consequence, evaluation is largely confined to factual retrieval, with little room for the deeper reasoning required for multi\-preference compounding or dynamic conflict resolution under real\-world constraints\.

To close this gap, we ground evaluation in authentic long\-term multimodal personal data rather than synthetic substrates\. Building on real, longitudinal records of participants’ daily lives—where every moment carries its own implicit personal context—we introduce the first benchmark constructed from genuine multimodal personal histories, together with a three\-tier task hierarchy that rigorously evaluates factual recall, preference and behavior inference, and complex predictive personalization\.

### 2\.2Egocentric and Lifelog Vision

Some research works study personal visual data directly\. Egocentric video understanding builds large first\-person datasets for activity and interaction recognition\[[10](https://arxiv.org/html/2609.19167#bib.bib32),[4](https://arxiv.org/html/2609.19167#bib.bib33),[11](https://arxiv.org/html/2609.19167#bib.bib39)\], and lifelog retrieval organizes and searches wearable\-camera or personal\-photo streams\[[12](https://arxiv.org/html/2609.19167#bib.bib34),[13](https://arxiv.org/html/2609.19167#bib.bib35)\]\. Recent efforts push toward longer horizons and assistant\-style use—extremely long egocentric video understanding\[[49](https://arxiv.org/html/2609.19167#bib.bib48)\]and egocentric life assistants\[[6](https://arxiv.org/html/2609.19167#bib.bib49)\]—while OmniQuery\[[25](https://arxiv.org/html/2609.19167#bib.bib19)\], closest to our setting, answers personal questions over captured multimodal memories\.

Yet these efforts largely operate on continuous egocentric video over hours to days, rely on third\-person or task labels rather than the data owner’s own subjective annotations, and stop short of the higher\-order persona inference and predictive personalization our tasks require\. ReaLMem advances this line from perception toward long\-term personalized cognition, pairing multi\-year personal visual archives with first\-person annotations and a graded, cognitively tiered evaluation\.

![Refer to caption](https://arxiv.org/html/2609.19167v1/dataset_construction.png)Figure 2:The ReaLMem construction pipeline: raw personal photos and videos are annotated by their owners \(objective and subjective layers\), anonymized, and woven into chronological multimodal sessions\.
### 2\.3Memory Mechanism for LLM

Endowing LLMs with long\-term memory has been pursued along two complementary directions\. The first scales raw access to history through ever\-longer context windows or retrieval\-augmented generation \(RAG\)\[[23](https://arxiv.org/html/2609.19167#bib.bib14),[41](https://arxiv.org/html/2609.19167#bib.bib15),[44](https://arxiv.org/html/2609.19167#bib.bib16)\], which fetches semantically similar historical snippets at query time\. The second introduces structured memory architectures that organize personal history into reusable representations, including knowledge\-graph memories\[[3](https://arxiv.org/html/2609.19167#bib.bib13),[35](https://arxiv.org/html/2609.19167#bib.bib12)\], hierarchical memory trees\[[26](https://arxiv.org/html/2609.19167#bib.bib9)\], and OS\-style memory frameworks that manage memory as a lifecycle\-controlled resource\[[27](https://arxiv.org/html/2609.19167#bib.bib10),[15](https://arxiv.org/html/2609.19167#bib.bib11)\]\.

Despite these advances, existing systems share two limitations directly relevant to long\-horizon personalization\. \(i\) Preference flattening: retrieval and storage treat all historical signals as equally valid, providing no principled mechanism to differentiate long\-held habits from one\-off mentions, so the model is left to arbitrate competing preferences on its own\. \(ii\) Weak temporal awareness: timestamps, when present, are typically used only as retrieval filters rather than as first\-class priors over content, leading to unresolved conflicts between outdated and recent records during reasoning\. Together, these gaps limit reliable behavior on tasks that require synthesizing heterogeneous, time\-varying signals into a coherent personalized decision\.

In response, ChronoProfiler builds a per\-user profile in which every item is grounded by timestamped evidence and weighted by a temporal stability score derived from its temporal features and used as a salience prior, enabling memory systems to resolve temporal conflicts and prioritize stable preferences without modifying the underlying LLM\.

## 3ReaLMem

We introduce ReaLMem \(Real\-world Long\-term Multimodal Memory\)\. Unlike prior benchmarks built from pre\-defined, synthesized personas, ReaLMem is constructed from authentic personal visual archives—photos and videos spanning years of real life—paired with first\-person subjective annotations from the data owners themselves\. This grounding in real multimodal experience and the owner’s own perspective provides a rigorous basis for studying memory retention, preference evolution, and personalized assistance\. We describe the dataset construction and task design below\.

### 3\.1Dataset Construction

We construct ReaLMem through a participant\-centric pipeline that turns raw personal visual archives into structured, high\-fidelity multimodal memory sequences \(Fig\.[2](https://arxiv.org/html/2609.19167#S2.F2)\)\. We recruited 7 participants with dense, multi\-year media collections, each contributing over 350 photos and videos that capture meaningful events, daily routines, and preference\-revealing behaviors \(*e\.g*\., dietary habits, long\-term hobbies\), all grounded with reliable spatio\-temporal metadata\. We deliberately favor depth over breadth: collecting truly authentic, long\-horizon archives is bounded by acquisition cost and strict privacy requirements, so a small set of richly annotated, real participants offers grounding that large synthetic or crowdsourced corpora cannot\. To ensure both factual accuracy and subjective authenticity, each participant followed a dual\-layer annotation protocol: an objective tier verified structured metadata, while a subjective tier recorded each record’s first\-person context—the personal event context, the owner’s feelings, and implicit preference signals\. All data collection, processing, and annotation followed an ethics protocol approved by our organization’s ethical review board, with written informed consent from every participant, details of ethical consideration are provided in Appendix[A](https://arxiv.org/html/2609.19167#A1)\.

After stringent anonymization: automated face blurring with RetinaFace\[[5](https://arxiv.org/html/2609.19167#bib.bib50)\], personally identifiable information \(PII\) pseudonymization, and metadata reorganization \(details in Appendix[B](https://arxiv.org/html/2609.19167#A2)\), we use GPT\-4o\[[16](https://arxiv.org/html/2609.19167#bib.bib36)\]to weave each participant’s records and annotations, in chronological order, into natural multi\-turn dialogues\. Crucially, the model only renders the narrative form: every fact, timestamp, and preference originates from the real records and the owner’s annotations, not from model invention\. As summarized in Tab\.[1](https://arxiv.org/html/2609.19167#S3.T1), the resulting data is diverse across time, space, and semantics: it spans 2018–2025, covers more than 40 countries and regions and ten everyday scenario categories, and comprises 2,333 images and 175 videos\. In total, ReaLMem contains 2,508 multimodal sessions \(about 358 per participant\), forming long\-term and coherent personal trajectories\.

Table 1:Statistics of the ReaLMem dataset\.TTdenotes the year of data collection\. Country counts reflect the top\-5 capturing locations by media volume\.Year \(rel\. toTT\)Media TypeTT857Image2,333T−1T\{\-\}1799Video175T−2T\{\-\}2327Country \(Top 5\)T−3T\{\-\}3285China941T−4T\{\-\}4148United Kingdom842T−5T\{\-\}558Japan127≤T−6\{\\leq\}\\,T\{\-\}634Spain114France74Scenario TypeLeisure & Entertain\.815Sports & Physical Act\.184Nature & Outdoor790Special Events175Home & Routine400Travel & Mobility145Social Interaction399Work & Study128Fashion27Medical & Care8
### 3\.2Task Design

Inspired by cognitive theories\[[1](https://arxiv.org/html/2609.19167#bib.bib24),[40](https://arxiv.org/html/2609.19167#bib.bib25)\]of human memory and reasoning, we design a hierarchical benchmark that reflects the progressive nature of human memory and cognition, framing personalized AI capability as a transition from information access to behavioral understanding and decision\-oriented application\. Specifically, we define three categories of increasing complexity, with per\-subtask definitions and example queries shown in Fig\.[1](https://arxiv.org/html/2609.19167#S1.F1)\(the full taxonomy with per\-subtask QA counts is given in Tab\.[7](https://arxiv.org/html/2609.19167#A4.T7)\)\.*Factual Memory Recall*\(T1 task\) tests foundational memory access grounded in evidence—retrieving event details \(task 1\.1\), aggregating or comparing across records \(task 1\.3\), and aligning information across the long\-term archive, including content available only visually, such as embedded text in visual data \(task 1\.2\) and memories grounded in a given photo \(task 1\.4\)\.*Persona Inference*\(T2 task\) moves from explicit observation to implicit patterns: profiling stable habits \(task 2\.1\), tracking how preferences evolve over time \(task 2\.2\), and reasoning about the motivations behind behavioral change \(task 2\.3\)\.*Predictive Personalization*\(T3 task\) turns this understanding into application: predicting users’ choices in novel scenarios \(task 3\.1\) and generating personalized recommendations based on personal history \(task 3\.2\)\. Unlike prior benchmarks based on synthetic personas or unimodal dialogue, this formulation is grounded in real multimodal experience and traces a continuous path from recall to decision\-making, enabling fine\-grained evaluation of long\-term personalization\.

QA Instantiation\.Each tier adopts a QA format aligned with its evaluation goal\. T1 questions are open\-ended, paired with a reference answer and explicit grounding evidence \(data IDs\)\. T2 questions augment the reference answer with a set of key supporting points that enable coverage evaluation of inherently subjective reasoning\. T3 adopts a ranking paradigm in which each query presents multiple preference\-constructed candidates, and the task is to recover the user’s actual preference ordering\. All 1,629 QA pairs are screened and refined by trained annotators\. Additionally, T2 and T3 instances are verified by the original data contributors to ensure faithfulness to lived and subjective experience\. Full construction procedures, quality\-control criteria, and ranking\-paradigm design are provided in Appendix[E](https://arxiv.org/html/2609.19167#A5)\.

## 4ChronoProfiler

Existing memory systems face two failure modes: \(i\) Temporal Conflicts, where outdated and recent signals coexist in context \(*e\.g*\., “vegan two years ago” vs\. “BBQ last week”\), confusing the model about the user’s current state; and \(ii\) Intra\-class Competition, where multiple same\-category preferences \(*e\.g*\., hotpot vs\. sushi\) compete without a principled arbitration mechanism\. Both issues are especially damaging on tasks that require understanding and reasoning over long\-term personal information, where the model must synthesize heterogeneous, time\-varying signals into a coherent judgment\. Inspired by memory consolidation\[[21](https://arxiv.org/html/2609.19167#bib.bib22),[30](https://arxiv.org/html/2609.19167#bib.bib23)\], in which repeatedly activated and recent experiences are reinforced while transient ones decay, we propose ChronoProfiler, which incrementally builds an increasingly comprehensive user profile as the history accumulates and assigns each user attribute item a temporal stability \(TS\) score that serves as a salience prior at inference\. The module runs in two stages, detailed below:

\(1\) Incremental Profile Construction\.We process the dialogue session\-by\-session, prompting an LLM to extract structured attributes into five semantic categories \(Demographics, Preferences, Behavioral Habits, Hobbies, Meaningful Facts\)\. Each candidate is semantically compared against existing entries in its category and resolved via one of four operations:New,Duplicate,Update, orConflict, yielding a comprehensive profile in which every item carries a set of timestamped evidence\. Implementation details are provided in Appendix[F](https://arxiv.org/html/2609.19167#A6)\.

\(2\) Temporal Stability Scoring\.For each profile itempp, let𝒮p\\mathcal\{S\}\_\{p\}denote the set of timestamps at whichpphas been mentioned across the dialogue history, and lettfirst=min⁡𝒮pt\_\{\\mathrm\{first\}\}=\\min\\mathcal\{S\}\_\{p\}andtlast=max⁡𝒮pt\_\{\\mathrm\{last\}\}=\\max\\mathcal\{S\}\_\{p\}be its earliest and latest mentions\. Given a reference timeTrefT\_\{\\mathrm\{ref\}\}\(set to the query timestamp at inference\), we summarize the lifecycle ofppwith three features:

ssup\\displaystyle s\_\{\\mathrm\{sup\}\}=1−e−\|𝒮p\|/τs,\\displaystyle=1\-e^\{\-\|\\mathcal\{S\}\_\{p\}\|/\\tau\_\{s\}\},sdur\\displaystyle s\_\{\\mathrm\{dur\}\}=1−e−\(tlast−tfirst\)/τd,\\displaystyle=1\-e^\{\-\(t\_\{\\mathrm\{last\}\}\-t\_\{\\mathrm\{first\}\}\)/\\tau\_\{d\}\},srec\\displaystyle s\_\{\\mathrm\{rec\}\}=e−\(Tref−tlast\)/τr\.\\displaystyle=e^\{\-\(T\_\{\\mathrm\{ref\}\}\-t\_\{\\mathrm\{last\}\}\)/\\tau\_\{r\}\}\.The Support scoressups\_\{\\mathrm\{sup\}\}measures*how often*ppis mentioned and saturates as\|𝒮p\|\|\\mathcal\{S\}\_\{p\}\|exceeds the characteristic countτs\\tau\_\{s\}\. The Duration scoresdurs\_\{\\mathrm\{dur\}\}measures*how long*the attribute persists, growing as the spantlast−tfirstt\_\{\\mathrm\{last\}\}\-t\_\{\\mathrm\{first\}\}exceeds the timescaleτd\\tau\_\{d\}\. The Recency scoresrecs\_\{\\mathrm\{rec\}\}measures*how recently*ppwas reaffirmed, decaying exponentially in the elapsed timeTref−tlastT\_\{\\mathrm\{ref\}\}\-t\_\{\\mathrm\{last\}\}with characteristic constantτr\\tau\_\{r\}\.

Each feature’s contribution adaptively per user via the entropy weight method\[[37](https://arxiv.org/html/2609.19167#bib.bib37),[50](https://arxiv.org/html/2609.19167#bib.bib38)\], a label\-free criterion\-weighting technique from multi\-criteria decision analysis\. Intuitively, a lifecycle feature that is nearly constant across a user’s profile items carries little discriminative information and should not drive prioritization, whereas a widely\-varying feature is informative and is up\-weighted\. Let𝐗∈ℝm×3\\mathbf\{X\}\\in\\mathbb\{R\}^\{m\\times 3\}collect the three feature scoressj​\(p\)∈\{ssup,sdur,srec\}s\_\{j\}\(p\)\\in\\\{s\_\{\\mathrm\{sup\}\},s\_\{\\mathrm\{dur\}\},s\_\{\\mathrm\{rec\}\}\\\}over the user’smmprofile items\. For each feature columnjjwe normalizepi​j=xi​j/∑ixi​jp\_\{ij\}=x\_\{ij\}/\\sum\_\{i\}x\_\{ij\}, compute its Shannon entropyej=−1ln⁡m∑ipi​jlnpi​je\_\{j\}=\-\\tfrac\{1\}\{\\ln m\}\\sum\_\{i\}p\_\{ij\}\\ln p\_\{ij\}, and weight it by the resulting information divergence:

wj=1−ej∑j′\(1−ej′\)\.w\_\{j\}=\\frac\{1\-e\_\{j\}\}\{\\sum\_\{j^\{\\prime\}\}\\left\(1\-e\_\{j^\{\\prime\}\}\\right\)\}\.The raw stability ofppis the entropy\-weighted sumzp=∑jwj​sj​\(p\)z\_\{p\}=\\sum\_\{j\}w\_\{j\}\\,s\_\{j\}\(p\)\. Finally, to obtain a salience prior comparable across users with heterogeneous engagement rhythms, we apply within\-user min–max normalization:

TS⁡\(p\)=zp−minq⁡zqmaxq⁡zq−minq⁡zq,\\mathrm\{TS\}\(p\)=\\frac\{z\_\{p\}\-\\min\_\{q\}z\_\{q\}\}\{\\max\_\{q\}z\_\{q\}\-\\min\_\{q\}z\_\{q\}\},spreading each user’s items across\[0,1\]\[0,1\]\. Long\-held and repeatedly reinforced preferences receive high TS \(consolidated long\-term memory\), whereas recent but unaccumulated signals receive moderate TS \(current salience without permanent commitment\)\. The TS\-scored profile is then injected into the LLM context as a salience prior, guiding the model to prioritize temporally stable signals during personalized decision\-making\. Parameter settings are provided in Appendix[F](https://arxiv.org/html/2609.19167#A6)\.

## 5Experiments

Table 2:Main results on theReaLMembenchmark \(%\), all metrics are scaled to\[0,100\]\[0,100\]\. Subtask abbreviations:Retr\./Embd\./Aggr\./M\.R\.= Episodic Event Retrieval / Embedded Content Retrieval / Multi\-Record Aggregation / Multimodal Recall;Prof\./Track/Reas\.= Persona Profiling / Longitudinal Tracking / Behavioral Reasoning;Pred\./Asst\.= Context\-Aware Prediction / Context\-Aware Assistance\. ShadedAvg\.columns report the task\-level mean\. Best per column*within each block*inbold\.### 5\.1Experimental Setup

We benchmark three categories of frontier approaches onReaLMem: \(1\)Full\-context MLLMs—Gemini\-3\.0\-Flash\[[9](https://arxiv.org/html/2609.19167#bib.bib47)\], GPT\-4\.1\-mini\[[33](https://arxiv.org/html/2609.19167#bib.bib44)\], and GPT\-5\.4\-mini\[[34](https://arxiv.org/html/2609.19167#bib.bib45)\]—receive the complete captioned history as a practical upper bound under unrestricted access\. \(2\)Memory systems—Mem0\[[3](https://arxiv.org/html/2609.19167#bib.bib13)\], MemOS\[[27](https://arxiv.org/html/2609.19167#bib.bib10)\], and EverMemOS\[[15](https://arxiv.org/html/2609.19167#bib.bib11)\]—compress per\-user histories into persistent structured memories and retrieve relevant content at query time; to isolate the memory architecture from the underlying LLM, all memory systems share a unified GPT\-4\.1\-mini answering backbone\. \(3\)ChronoProfiler \(Ours\)performs a controlled*profile swap*: holding EverMemOS episodic retrieval and the reader backbone fixed, we replace EverMemOS’s built\-in user profile with our RAG\-retrieved temporally\-weighted profile \(§[4](https://arxiv.org/html/2609.19167#S4)\), and evaluate the same three MLLMs as readers on T2 and T3 tasks \(the profile does not affect T1 factual recall\)\. As an additional T1 upper bound \(Tab\.[3](https://arxiv.org/html/2609.19167#S5.T3)\), we conductOracleexperiments that supply the three frontier MLLMs with only the gold\-evidence sessions \(identified by per\-QA annotated evidence id\) as context\. Unless otherwise specified, all evaluations operate on a unified text representation in which media files are pre\-captioned by Gemini\-2\.0\-Flash\[[8](https://arxiv.org/html/2609.19167#bib.bib46)\]and embedded into the dialogue transcript\. Full configuration details and ChronoProfiler hyperparameters are in Appendices[C](https://arxiv.org/html/2609.19167#A3)and[F](https://arxiv.org/html/2609.19167#A6)\.

Evaluation Metrics\.We use GPT\-4o\-mini\[[31](https://arxiv.org/html/2609.19167#bib.bib43)\]as an LLM judge\[[47](https://arxiv.org/html/2609.19167#bib.bib40)\]forT1task \(binary correctness against the reference answer\), the normalized mean ofCoverageandAccuracyscore forT2task, and Kendall\-τ\\taurank correlation\[[20](https://arxiv.org/html/2609.19167#bib.bib41)\]against the participant\-validated ground\-truth ordering forT3task\. All metrics are scaled to\[0,100\]\[0,100\]; scoring rubrics and evaluation details are provided in Appendix[D](https://arxiv.org/html/2609.19167#A4)\.

### 5\.2Main Results

Tab\.[2](https://arxiv.org/html/2609.19167#S5.T2)presents results across all evaluated MLLMs and systems, and Tab\.[3](https://arxiv.org/html/2609.19167#S5.T3)reports the Oracle T1 upper bound\.

Table 3:Oracle reference on T1 Factual Memory Recall \(%\)\. For each model the top row reports the Oracle score, the bottom row reports the absolute delta versus the same model’s full\-context result in Tab\.[2](https://arxiv.org/html/2609.19167#S5.T2)\.Full\-context performance confirms the cognitive hierarchy\.T3 is the uniformly hardest tier across all three frontier MLLMs \(57\.9–59\.2%\), confirming a consistent upper bound on predictive\-personalization difficulty\. Above T3, the T1–T2 performance is model\-dependent: Gemini\-3\.0\-Flash follows the expected cognitive hierarchy \(T1 82\.6%\>\>T2 76\.4%\), whereas GPT\-4\.1\-mini and GPT\-5\.4\-mini both score marginally higher on T2 than T1, indicating that preference inference is not universally harder than factual recall and that long\-context retrieval ability is the primary T1 differentiator\. Gemini\-3\.0\-Flash leads on both T1 \(82\.6%\) and T2 \(76\.4%\), its T1 advantage is concentrated on Multi\-Record Aggregation and Multimodal Recall\. On T2 and T3, the three models converge tightly, spanning 2\.1 points on T2 and 1\.3 points on T3, indicating that difficulty is dominated by the inherent complexity of preference reasoning and predictive personalization rather than by backbone capacity\.

Table 4:Effect of context modality on Gemini\-3\.0\-Flash \(%\)\.Cap\.\+Dial\.: captioned dialogue with metadata;MM: raw visual media with dialogue and metadata;Vis\.\+Meta\.: visual media with metadata only \(no dialogue\)\. Task abbreviations follow Tab\.[2](https://arxiv.org/html/2609.19167#S5.T2)\. Best result per row inbold\.Memory architectures shape the compression–quality profile\.Under the unified GPT\-4\.1\-mini answering backbone \(full\-context T1/T2/T3 average==72\.6/74\.3/57\.9%\), the three memory systems exhibit markedly different compression–quality trade\-offs at∼70×\{\\sim\}70\\timescompression that map cleanly onto their architectural choices\. Mem0\[[3](https://arxiv.org/html/2609.19167#bib.bib13)\], which extracts flat narrative memories via LLM\-drivenADD/UPDATEoperations, discards precise temporal and quantitative anchors and suffers the largest T1 collapse \(T1 46\.8%\)\. MemOS\[[27](https://arxiv.org/html/2609.19167#bib.bib10)\], with OS\-style hierarchical MemCube governance, recovers more structured factual recall \(T1 54\.0%\) and attaches a static global preference summary to each retrieved memory; however, this uniformly\-replicated, coarsely\-aggregated preference string provides insufficient query\-specific grounding, capping T2 at 58\.8%\. EverMemOS\[[15](https://arxiv.org/html/2609.19167#bib.bib11)\], whose engram\-inspired Episodic Trace and Semantic Consolidation lifecycle distils both events and preferences, tracks the full\-context baseline most closely \(T1 69\.4%\) and attains the strongest memory\-system T2 \(72\.0%\)\. Strikingly, on T3 all three systems land within about one point of the GPT\-4\.1\-mini full\-context baseline \(56\.8–59\.2% vs\. 57\.9%\), with Mem0 even edging ahead \(59\.2%\), indicating that the T3 bottleneck is representational quality rather than memory coverage\.

Table 5:Profile\-swap analysis on T2 and T3 \(%\)\.ChronoProfilerinjects our temporally\-weighted profile,Ep\.\+Prof\.EverMemOS’s built\-in profile; for each row, the top line is the score and the bottom line itsΔ\\Deltaversus the same backbone’s Full Context \(Tab\.[2](https://arxiv.org/html/2609.19167#S5.T2)\)\. Best per column inbold\.Oracle exposes a retrieval bottleneck on T1\.All three MLLMs converge to a similar Oracle ceiling \(85\.4–85\.6% T1 Avg\.\), but their full\-context gaps diverge sharply: Gemini trails its Oracle by only\+3\.0%\+3\.0\\%, whereas GPT\-4\.1\-mini and GPT\-5\.4\-mini both lag far more \(\+12\.9%\+12\.9\\%each\)\. This shows that Gemini’s T1 advantage stems primarily from robustness to distractor evidence in long context, and memory systems implicitly inherit the task of approximating an oracle retriever—which is why EverMemOS can close most of the GPT\-4\.1\-mini full\-context gap with only 1\.7k tokens\.

T3 reveals a universal ceiling unaddressed by context length or retrieval\.Across every system in Tab\.[2](https://arxiv.org/html/2609.19167#S5.T2), T3 remains the hardest level, with the highest T3 Avg\. capped at 59\.2% and the full spread under 3 points \(56\.8–59\.2%\)\. Unlike T1, T3 has no evidence\-bounded upper bound: the bottleneck is not retrieval or compression but the fine\-grained discrimination of preference signals differentially influence decisions across contexts—recognizing that a user holds a preference is necessary but insufficient; the model must further judge which preferences dominate, when they trade off, and how their relative weights shift with situational context\.

### 5\.3What Each Modality Contributes

To probe the effect of context modality, we compare three input settings: captioned dialogue with spatio\-temporal metadata \(our default\), raw multimodal media with dialogue and metadata \(MM\), and visual media with metadata but no dialogue \(Vis\.\+Meta\.\)\. Because raw multimodal input far exceeds the captioned context length, we run this ablation on Gemini\-3\.0\-Flash, the reader that natively ingests long\-context raw multimodal input \(Tab\.[4](https://arxiv.org/html/2609.19167#S5.T4)\)\.

Captions are a faithful proxy for raw pixels on factual recall\.With dialogue held fixed, captioned dialogue and MM differ only marginally—captions slightly help text\-anchored subtasks \(episodic and embedded retrieval, all of T2\), while raw pixels slightly help the two subtasks needing quantitative or visual reasoning \(Multi\-Record Aggregation and Multimodal Recall\)\. The overall T1 gap is just1\.61\.6points \(82\.6 vs\. 81\.0\), showing that high\-quality MLLM captions preserve most factual content and validating captioned dialogue as the unified input for fair comparison across memory systems\.

Language\-mediated context dominates recall and inference tasks\.Holding the visual input fixed and removing only the dialogue stream \(MM→\\toVis\.\+Meta\.\) collapses T1 by23\.623\.6points \(81\.0→\\to57\.4\) and T2 by12\.412\.4\(74\.8→\\to62\.4\), sharpest on Episodic Event Retrieval \(−32\.1\-32\.1\)\. The drop is this steep because most of ReaLMem’s factual grounding, together with all of its first\-person subjective annotations resides in the textual stream: raw frames depict scenes but not the identities, labels, and personal context that T1 recall and T2 inference require\.

Vision alone is the strongest signal for prediction\.T3 breaks the monotone pattern: the dialogue\-free setting \(Vis\.\+Meta\.\) attains the highest T3 Avg\. \(61\.461\.4vs\.59\.259\.2\) and leads on both subtasks\. We attribute this gap to the nature of the signal: the visual stream itself reflects a user’s preferences and the relative composition of their life\. A rich, multi\-year dialogue, by contrast, injects mutually conflicting and uneven signals that the model cannot reliably arbitrate, whereas the sparser visual\-and\-metadata view falls back on these broad, stable patterns of activity and place\. Since predictive personalization is a ranking task that rewards a holistic read of a user’s habitual preferences over fine\-grained arbitration of scattered textual cues, this sparser view wins\.

### 5\.4ChronoProfiler Analysis

We probe ChronoProfiler with two controlled experiments, both evaluating the three MLLMs on T2 and T3 with deltas reported against the same backbone’s Full Context\. First, a*profile swap*\(Tab\.[5](https://arxiv.org/html/2609.19167#S5.T5)\) holds EverMemOS episodic retrieval and replaces EverMemOS’s built\-in profile with our temporally\-weighted profile; comparing against Full Context and againstEverMemOS: Ep\.\+Prof\.\(the identical pipeline with EverMemOS’s own profile\) isolates the profile representation from raw context length and from the quality of a profile\. Second, a*TS ablation*\(Tab\.[6](https://arxiv.org/html/2609.19167#S5.T6)\) fixes the retrieval depth \(top\-15 items retrieval per category\) and compares the profile with and without temporal stability weighting, isolating the contribution of TS\. We additionally study how the retrieval depth affects performance, the configuration and full depth–accuracy analysis are deferred to Appendix[G](https://arxiv.org/html/2609.19167#A7), with all main experiments fixing the depth at top\-15\. Together, these analyses separate the factors behind the gains: context length, profile representation, TS weighting, and retrieval depth\.

ChronoProfiler matches or beats Full Context on both T2 and T3\.At only∼6\{\\sim\}6k prompt tokens, roughly20×20\\timesfewer than the Full Context, ChronoProfiler exceeds full context on T2 Avg\. for all three readers \(Δ=\+0\.6/\+4\.0/\+4\.3\\Delta=\+0\.6/\+4\.0/\+4\.3\) and matches or exceeds it on T3 Avg\. \(Δ=±0\.0/\+3\.5/\+1\.8\\Delta=\\pm 0\.0/\+3\.5/\+1\.8\)\. The strongest configuration is GPT\-4\.1\-mini, the only reader to improve on*every*subtask, posting the table\-best T3 Avg\. \(61\.4%,Δ=\+3\.5\\Delta=\+3\.5\) and the best Context\-Aware Assistance \(57\.8%\)\. This confirms that the context sufficient for preference\-grounded inference is far smaller than the raw history, provided it captures the right temporal structure\.

The temporally\-weighted profile dominates EverMemOS’s built\-in profile\.Holding episodes and the GPT\-4\.1\-mini backbone fixed, swapping in our profile improves*every*T2 metric \(T2 Avg\.\+6\.3\+6\.3; Persona Profiling\+5\.6\+5\.6, Longitudinal Tracking\+6\.9\+6\.9, Behavioral Reasoning\+6\.8\+6\.8\) and both T3 metrics \(T3 Avg\.\+2\.1\+2\.1; Context\-Aware Assistance\+4\.3\+4\.3\)\. Tellingly, EverMemOS’s own profile*degrades*Full Context on T2 \(Δ=−2\.3\\Delta=\-2\.3\) while ours*lifts*it \(Δ=\+4\.0\\Delta=\+4\.0\), pinpointing the profile representation—not the episodic memory—as the active ingredient behind the T2 and T3 gains\.

Temporal stability is the source of the T2 gains\.Adding TS lifts T2 Avg\. for all three readers \(\+2\.7/\+2\.0/\+0\.3\+2\.7/\+2\.0/\+0\.3\) and raises Persona Profiling across the board \(\+3\.0/\+2\.6/\+1\.1\+3\.0/\+2\.6/\+1\.1\), directly evidencing TS as a time\-aware prior for consolidating stable and evolving preferences\. On T3 the effect is comparable or slightly lower \(T3 Avg\+0\.6/−0\.4/−1\.2\+0\.6/\-0\.4/\-1\.2\): TS consistently helps Context\-Aware Prediction for both GPT readers \(\+2\.4/\+0\.9\+2\.4/\+0\.9\) but uniformly trades off a small amount of Context\-Aware Assistance \(≈−2\.4\{\\approx\}\-2\.4\), which we attribute to TS sharpening recurring preferences at the expense of the broader, occasion\-specific cues that open\-ended recommendation can exploit\. We therefore scope the TS claim to persona profiling and longitudinal tracking, and regard selective episodic retrieval as a complementary route to recovering the Assistance gap\.

Table 6:Temporal\-stability \(TS\) ablation on T2 and T3 \(%\)\. Rows compare the profile without \(*w/o TS*\) and with \(\+ TS\) TS weighting;Δ\\Deltais the gain from adding TS\.

## 6Conclusion and Limitations

We introduced ReaLMem, the first multimodal long\-term memory benchmark grounded in authentic personal visual archives and first\-person subjective annotations, organized as a Bloom\-inspired three\-tier hierarchy from factual recall through persona inference to predictive personalization\. To counter the preference\-flattening of existing memory systems, we proposed ChronoProfiler, a memory\-consolidation\-inspired module that scores each attribute by its temporal stability and injects it as a salience prior\. Our experiments identify predictive personalization as the central bottleneck and show that a compact temporally\-weighted profile can rival full context, establishing temporally\-weighted preference selection as a promising direction for real\-world personal\-memory modeling\.

ReaLMem deliberately trades breadth for depth: authentic, long\-horizon first\-person archives are bounded by acquisition cost and strict privacy, so the participant pool is small, and ChronoProfiler’s decay constants follow cognitive principles rather than population\-scale tuning\. As larger and more diverse personal\-memory corpora emerge, broader\-coverage benchmarks and adaptive, user\-specific temporal\-stability parameterizations—learned from each user’s own engagement rhythm—are natural next steps\.

## References

- \[1\]B\. S\. Bloom, M\. D\. Engelhart, E\. J\. Furst, W\. H\. Hill, D\. R\. Krathwohl,et al\.\(1956\)Handbook i: cognitive domain\.New York: David McKay,pp\. 483–498\.Cited by:[§1](https://arxiv.org/html/2609.19167#S1.p6.1),[§3\.2](https://arxiv.org/html/2609.19167#S3.SS2.p1.1)\.
- \[2\]T\. F\. Brady, T\. Konkle, G\. A\. Alvarez, and A\. Oliva\(2008\)Visual long\-term memory has a massive storage capacity for object details\.Proceedings of the National Academy of Sciences105,pp\. 14325 – 14329\.External Links:[Link](https://api.semanticscholar.org/CorpusID:1211873)Cited by:[§1](https://arxiv.org/html/2609.19167#S1.p2.1)\.
- \[3\]P\. Chhikara, D\. Khant, S\. Aryan, T\. Singh, and D\. Yadav\(2025\)Mem0: building production\-ready AI agents with scalable long\-term memory\.InECAI 2025 \- 28th European Conference on Artificial Intelligence, 25\-30 October 2025, Bologna, Italy \- Including 14th Conference on Prestigious Applications of Intelligent Systems \(PAIS 2025\),I\. Lynce, N\. Murano, M\. Vallati, S\. Villata, F\. Chesani, M\. Milano, A\. Omicini, and M\. Dastani \(Eds\.\),Frontiers in Artificial Intelligence and Applications, Vol\.413,pp\. 2993–3000\.External Links:[Link](https://doi.org/10.3233/FAIA251160),[Document](https://dx.doi.org/10.3233/FAIA251160)Cited by:[Appendix C](https://arxiv.org/html/2609.19167#A3.SS0.SSS0.Px1.p1.1),[§2\.3](https://arxiv.org/html/2609.19167#S2.SS3.p1.1),[§5\.1](https://arxiv.org/html/2609.19167#S5.SS1.p1.1),[§5\.2](https://arxiv.org/html/2609.19167#S5.SS2.p3.1)\.
- \[4\]D\. Damen, H\. Doughty, G\. M\. Farinella, A\. Furnari, E\. Kazakos, J\. Ma, D\. Moltisanti, J\. Munro, T\. Perrett, W\. Price, and M\. Wray\(2020\)Rescaling egocentric vision: collection, pipeline and challenges for epic\-kitchens\-100\.International Journal of Computer Vision130,pp\. 33 – 55\.External Links:[Link](https://api.semanticscholar.org/CorpusID:244619848)Cited by:[§1](https://arxiv.org/html/2609.19167#S1.p5.1),[§2\.2](https://arxiv.org/html/2609.19167#S2.SS2.p1.1)\.
- \[5\]J\. Deng, J\. Guo, E\. Ververas, I\. Kotsia, and S\. Zafeiriou\(2020\)RetinaFace: single\-shot multi\-level face localisation in the wild\.InCVPR,Cited by:[Appendix B](https://arxiv.org/html/2609.19167#A2.p1.1),[§3\.1](https://arxiv.org/html/2609.19167#S3.SS1.p2.1)\.
- \[6\]Y\. Dong and Y\. Dong\(2025\)EgoLife: towards egocentric life assistant\.2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 28885–28900\.External Links:[Link](https://api.semanticscholar.org/CorpusID:276883706)Cited by:[§1](https://arxiv.org/html/2609.19167#S1.p5.1),[§2\.2](https://arxiv.org/html/2609.19167#S2.SS2.p1.1)\.
- \[7\]Y\. Du, H\. Wang, Z\. Zhao, B\. Liang, B\. Wang, W\. Zhong, Z\. Wang, and K\. Wong\(2024\)PerLTQA: a personal long\-term memory dataset for memory classification, retrieval, and fusion in question answering\.InProceedings of the 10th SIGHAN Workshop on Chinese Language Processing \(SIGHAN\-10\),pp\. 152–164\.Cited by:[§1](https://arxiv.org/html/2609.19167#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.19167#S2.SS1.p1.1)\.
- \[8\]Google DeepMind\(2025\)Gemini 2\.0: flash, flash\-lite and pro\.Note:[https://blog\.google/technology/google\-deepmind/gemini\-model\-updates\-february\-2025/](https://blog.google/technology/google-deepmind/gemini-model-updates-february-2025/)Gemini\-2\.0\-Flash; accessed 2026\-06\-26Cited by:[§5\.1](https://arxiv.org/html/2609.19167#S5.SS1.p1.1)\.
- \[9\]Google DeepMind\(2026\)Introducing gemini 3 flash\.Note:[https://blog\.google/products\-and\-platforms/products/gemini/gemini\-3\-flash/](https://blog.google/products-and-platforms/products/gemini/gemini-3-flash/)Gemini\-3\.0\-Flash; accessed 2026\-06\-26Cited by:[§5\.1](https://arxiv.org/html/2609.19167#S5.SS1.p1.1)\.
- \[10\]K\. Graumanet al\.\(2021\)Ego4D: around the world in 3,000 hours of egocentric video\.2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 18973–18990\.External Links:[Link](https://api.semanticscholar.org/CorpusID:238856888)Cited by:[§1](https://arxiv.org/html/2609.19167#S1.p5.1),[§2\.2](https://arxiv.org/html/2609.19167#S2.SS2.p1.1)\.
- \[11\]K\. Graumanet al\.\(2023\)Ego\-exo4d: understanding skilled human activity from first\- and third\-person perspectives\.2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 19383–19400\.External Links:[Link](https://api.semanticscholar.org/CorpusID:265506384)Cited by:[§2\.2](https://arxiv.org/html/2609.19167#S2.SS2.p1.1)\.
- \[12\]C\. Gurrin, H\. Joho, F\. Hopfgartner, L\. Zhou, and R\. Albatal\(2016\)Overview of ntcir\-12 lifelog task\.InNTCIR Conference on Evaluation of Information Access Technologies,External Links:[Link](https://api.semanticscholar.org/CorpusID:10169871)Cited by:[§2\.2](https://arxiv.org/html/2609.19167#S2.SS2.p1.1)\.
- \[13\]C\. Gurrin, A\. F\. Smeaton, and A\. R\. Doherty\(2014\)LifeLogging: personal big data\.Found\. Trends Inf\. Retr\.8,pp\. 1–125\.External Links:[Link](https://api.semanticscholar.org/CorpusID:62586218)Cited by:[§2\.2](https://arxiv.org/html/2609.19167#S2.SS2.p1.1)\.
- \[14\]C\. Hsieh, S\. Sun, S\. Kriman, S\. Acharya, D\. Rekesh, F\. Jia, Y\. Zhang, and B\. Ginsburg\(2024\)RULER: what’s the real context size of your long\-context language models?\.CoRRabs/2404\.06654\.External Links:[Link](https://doi.org/10.48550/arXiv.2404.06654),[Document](https://dx.doi.org/10.48550/ARXIV.2404.06654),2404\.06654Cited by:[§1](https://arxiv.org/html/2609.19167#S1.p2.1)\.
- \[15\]C\. Hu, X\. Gao, Z\. Zhou, D\. Xu, Y\. Bai, X\. Li, H\. Zhang, T\. Li, C\. Zhang, L\. Bing,et al\.\(2026\)EverMemOS: a self\-organizing memory operating system for structured long\-horizon reasoning\.arXiv preprint arXiv:2601\.02163\.Cited by:[Appendix C](https://arxiv.org/html/2609.19167#A3.SS0.SSS0.Px1.p1.1),[§2\.3](https://arxiv.org/html/2609.19167#S2.SS3.p1.1),[§5\.1](https://arxiv.org/html/2609.19167#S5.SS1.p1.1),[§5\.2](https://arxiv.org/html/2609.19167#S5.SS2.p3.1)\.
- \[16\]A\. Hurstet al\.\(2024\)GPT\-4o system card\.External Links:[Link](https://api.semanticscholar.org/CorpusID:273662196)Cited by:[§3\.1](https://arxiv.org/html/2609.19167#S3.SS1.p2.1)\.
- \[17\]J\. Jang, M\. Boo, and H\. Kim\(2023\)Conversation chronicles: towards diverse temporal and relational dynamics in multi\-session conversations\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 13584–13606\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.838/),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.838)Cited by:[§2\.1](https://arxiv.org/html/2609.19167#S2.SS1.p1.1)\.
- \[18\]B\. Jiang, Z\. Hao, Y\. Cho, B\. Li, Y\. Yuan, S\. Chen, L\. Ungar, C\. J\. Taylor, and D\. Roth\(2025\)Know me, respond to me: benchmarking LLMs for dynamic user profiling and personalized responses at scale\.arXiv preprint arXiv:2504\.14225\.Cited by:[§1](https://arxiv.org/html/2609.19167#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.19167#S2.SS1.p1.1)\.
- \[19\]B\. Jiang, Y\. Yuan, M\. Shen, Z\. Hao, Z\. Xu, Z\. Chen, Z\. Liu, A\. R\. Vijjini, J\. He, H\. Yu, R\. Poovendran, G\. W\. Wornell, L\. H\. Ungar, D\. Roth, S\. Chen, and C\. J\. Taylor\(2025\)PersonaMem\-v2: towards personalized intelligence via learning implicit user personas and agentic memory\.CoRRabs/2512\.06688\.External Links:[Link](https://doi.org/10.48550/arXiv.2512.06688),[Document](https://dx.doi.org/10.48550/ARXIV.2512.06688),2512\.06688Cited by:[§1](https://arxiv.org/html/2609.19167#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.19167#S2.SS1.p1.1)\.
- \[20\]M\. G\. Kendall\(1938\)A new measure of rank correlation\.Biometrika30,pp\. 81–93\.External Links:[Link](https://api.semanticscholar.org/CorpusID:120478295)Cited by:[§5\.1](https://arxiv.org/html/2609.19167#S5.SS1.p2.1)\.
- \[21\]J\. G\. Klinzing, N\. Niethard, and J\. Born\(2019\)Mechanisms of systems memory consolidation during sleep\.Nature neuroscience22\(10\),pp\. 1598–1610\.Cited by:[§4](https://arxiv.org/html/2609.19167#S4.p1.1)\.
- \[22\]E\. Lee\(2024\)Towards ethical personal AI applications: practical considerations for AI assistants with long\-term memory\.arXiv preprint arXiv:2409\.11192\.Cited by:[§1](https://arxiv.org/html/2609.19167#S1.p1.1)\.
- \[23\]P\. Lewis, E\. Perez, A\. Piktus, F\. Petroni, V\. Karpukhin, N\. Goyal, H\. Küttler, M\. Lewis, W\. Yih, T\. Rocktäschel,et al\.\(2020\)Retrieval\-augmented generation for knowledge\-intensive NLP tasks\.Advances in neural information processing systems33,pp\. 9459–9474\.Cited by:[§2\.3](https://arxiv.org/html/2609.19167#S2.SS3.p1.1)\.
- \[24\]H\. Li, C\. Yang, A\. Zhang, Y\. Deng, X\. Wang, and T\. Chua\(2025\)Hello again\! LLM\-powered personalized agent for long\-term dialogue\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 5259–5276\.Cited by:[§2\.1](https://arxiv.org/html/2609.19167#S2.SS1.p1.1)\.
- \[25\]J\. N\. Li, Z\. Zhang, and J\. Ma\(2025\)OmniQuery: contextually augmenting captured multimodal memories to enable personal question answering\.InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems,pp\. 1–20\.Cited by:[§2\.2](https://arxiv.org/html/2609.19167#S2.SS2.p1.1)\.
- \[26\]X\. Li, J\. Bantupalli, R\. Dharmani, Y\. Zhang, and J\. Shang\(2025\)Toward multi\-session personalized conversation: a large\-scale dataset and hierarchical tree framework for implicit reasoning\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 11504–11517\.Cited by:[§2\.3](https://arxiv.org/html/2609.19167#S2.SS3.p1.1)\.
- \[27\]Z\. Li, S\. Song, H\. Wang, S\. Niu, D\. Chen, J\. Yang, C\. Xi, H\. Lai, J\. Zhao, Y\. Wang, J\. Ren, Z\. Lin, J\. Huo, T\. Chen, K\. Chen, K\. Li, Z\. Yin, Q\. Yu, B\. Tang, H\. Yang, Z\. J\. Xu, and F\. Xiong\(2025\)MemOS: an operating system for memory\-augmented generation \(MAG\) in large language models\.CoRRabs/2505\.22101\.External Links:[Link](https://doi.org/10.48550/arXiv.2505.22101),[Document](https://dx.doi.org/10.48550/ARXIV.2505.22101),2505\.22101Cited by:[Appendix C](https://arxiv.org/html/2609.19167#A3.SS0.SSS0.Px1.p1.1),[§2\.3](https://arxiv.org/html/2609.19167#S2.SS3.p1.1),[§5\.1](https://arxiv.org/html/2609.19167#S5.SS1.p1.1),[§5\.2](https://arxiv.org/html/2609.19167#S5.SS2.p3.1)\.
- \[28\]N\. F\. Liu, K\. Lin, J\. Hewitt, A\. Paranjape, M\. Bevilacqua, F\. Petroni, and P\. Liang\(2024\)Lost in the middle: how language models use long contexts\.Transactions of the association for computational linguistics12,pp\. 157–173\.Cited by:[§1](https://arxiv.org/html/2609.19167#S1.p7.1)\.
- \[29\]A\. Maharana, D\. Lee, S\. Tulyakov, M\. Bansal, F\. Barbieri, and Y\. Fang\(2024\)Evaluating very long\-term conversational memory of LLM agents\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 13851–13870\.Cited by:[§1](https://arxiv.org/html/2609.19167#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.19167#S2.SS1.p1.1)\.
- \[30\]J\. L\. McClelland, B\. L\. McNaughton, and R\. C\. O’Reilly\(1995\)Why there are complementary learning systems in the hippocampus and neocortex: insights from the successes and failures of connectionist models of learning and memory\.\.Psychological review102\(3\),pp\. 419\.Cited by:[§4](https://arxiv.org/html/2609.19167#S4.p1.1)\.
- \[31\]OpenAI\(2024\)GPT\-4o mini: advancing cost\-efficient intelligence\.Note:[https://openai\.com/index/gpt\-4o\-mini\-advancing\-cost\-efficient\-intelligence/](https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/)accessed 2026\-06\-26Cited by:[§5\.1](https://arxiv.org/html/2609.19167#S5.SS1.p2.1)\.
- \[32\]OpenAI\(2024\)New embedding models and API updates\.Note:[https://openai\.com/index/new\-embedding\-models\-and\-api\-updates/](https://openai.com/index/new-embedding-models-and-api-updates/)text\-embedding\-3\-small; accessed 2026\-06\-26Cited by:[Appendix F](https://arxiv.org/html/2609.19167#A6.SS0.SSS0.Px2.p1.1)\.
- \[33\]OpenAI\(2025\)Introducing GPT\-4\.1 in the API\.Note:[https://openai\.com/index/gpt\-4\-1/](https://openai.com/index/gpt-4-1/)GPT\-4\.1\-mini; accessed 2026\-06\-26Cited by:[§5\.1](https://arxiv.org/html/2609.19167#S5.SS1.p1.1)\.
- \[34\]OpenAI\(2025\)Introducing GPT\-5\.Note:[https://openai\.com/index/introducing\-gpt\-5/](https://openai.com/index/introducing-gpt-5/)GPT\-5\.4\-mini point release; replace with the 5\.4 model card if available; accessed 2026\-06\-26Cited by:[§5\.1](https://arxiv.org/html/2609.19167#S5.SS1.p1.1)\.
- \[35\]P\. Rasmussen, P\. Paliychuk, T\. Beauvais, J\. Ryan, and D\. Chalef\(2025\)Zep: A temporal knowledge graph architecture for agent memory\.CoRRabs/2501\.13956\.External Links:[Link](https://doi.org/10.48550/arXiv.2501.13956),[Document](https://dx.doi.org/10.48550/ARXIV.2501.13956),2501\.13956Cited by:[§2\.3](https://arxiv.org/html/2609.19167#S2.SS3.p1.1)\.
- \[36\]T\. Schuster, M\. Lambert, N\. Döring, and J\. Trögele\(2025\)Needle\-in\-the\-haystack testing LLMs with a complex reasoning task\.InInternational Conference on Engineering Applications of Neural Networks,pp\. 254–266\.Cited by:[§1](https://arxiv.org/html/2609.19167#S1.p2.1)\.
- \[37\]C\. E\. Shannon\(1948\)A mathematical theory of communication\.Bell Syst\. Tech\. J\.27,pp\. 623–656\.External Links:[Link](https://api.semanticscholar.org/CorpusID:55379485)Cited by:[§4](https://arxiv.org/html/2609.19167#S4.p4.1)\.
- \[38\]H\. Sun, Z\. Zhang, and S\. Zeng\(2025\)Preference\-aware memory update for long\-term LLM agents\.arXiv preprint arXiv:2510\.09720\.Cited by:[§1](https://arxiv.org/html/2609.19167#S1.p7.1)\.
- \[39\]M\. Tavakoli, A\. Salemi, C\. Ye, M\. Abdalla, H\. Zamani, and J\. R\. Mitchell\(2025\)Beyond a million tokens: benchmarking and enhancing long\-term memory in LLMs\.arXiv preprint arXiv:2510\.27246\.Cited by:[§1](https://arxiv.org/html/2609.19167#S1.p1.1)\.
- \[40\]E\. Tulving\(1972\)Episodic and semantic memory\.External Links:[Link](https://api.semanticscholar.org/CorpusID:140830322)Cited by:[§3\.2](https://arxiv.org/html/2609.19167#S3.SS2.p1.1)\.
- \[41\]W\. Wang, L\. Dong, H\. Cheng, X\. Liu, X\. Yan, J\. Gao, and F\. Wei\(2023\)Augmenting language models with long\-term memory\.Advances in Neural Information Processing Systems36,pp\. 74530–74543\.Cited by:[§2\.3](https://arxiv.org/html/2609.19167#S2.SS3.p1.1)\.
- \[42\]R\. Westhäußer, W\. Minker, and S\. Zepf\(2025\)Enabling personalized long\-term interactions in LLM\-based agents through persistent memory and user profiles\.CoRRabs/2510\.07925\.External Links:[Link](https://doi.org/10.48550/arXiv.2510.07925),[Document](https://dx.doi.org/10.48550/ARXIV.2510.07925),2510\.07925Cited by:[§1](https://arxiv.org/html/2609.19167#S1.p7.1)\.
- \[43\]D\. Wu, H\. Wang, W\. Yu, Y\. Zhang, K\. Chang, and D\. Yu\(2025\)LongMemEval: benchmarking chat assistants on long\-term interactive memory\.InThe Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24\-28, 2025,External Links:[Link](https://openreview.net/forum?id=pZiyCaVuti)Cited by:[§1](https://arxiv.org/html/2609.19167#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.19167#S2.SS1.p1.1)\.
- \[44\]G\. Xiao, Y\. Tian, B\. Chen, S\. Han, and M\. Lewis\(2024\)Efficient streaming language models with attention sinks\.InThe Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7\-11, 2024,External Links:[Link](https://openreview.net/forum?id=NG7sS51zVF)Cited by:[§2\.3](https://arxiv.org/html/2609.19167#S2.SS3.p1.1)\.
- \[45\]X\. Xu, Z\. Gou, W\. Wu, Z\. Niu, H\. Wu, H\. Wang, and S\. Wang\(2022\)Long time no see\! open\-domain conversation with long\-term persona memory\.InFindings of the Association for Computational Linguistics: ACL 2022,pp\. 2639–2650\.Cited by:[§2\.1](https://arxiv.org/html/2609.19167#S2.SS1.p1.1)\.
- \[46\]Z\. Zhang, Q\. Dai, X\. Bo, C\. Ma, R\. Li, X\. Chen, J\. Zhu, Z\. Dong, and J\. Wen\(2025\)A survey on the memory mechanism of large language model\-based agents\.ACM Transactions on Information Systems43\(6\),pp\. 1–47\.Cited by:[§1](https://arxiv.org/html/2609.19167#S1.p1.1)\.
- \[47\]L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. P\. Xing, H\. Zhang, J\. E\. Gonzalez, and I\. Stoica\(2023\)Judging llm\-as\-a\-judge with mt\-bench and chatbot arena\.ArXivabs/2306\.05685\.External Links:[Link](https://api.semanticscholar.org/CorpusID:259129398)Cited by:[§5\.1](https://arxiv.org/html/2609.19167#S5.SS1.p2.1)\.
- \[48\]W\. Zhong, L\. Guo, Q\. Gao, H\. Ye, and Y\. Wang\(2024\)MemoryBank: enhancing large language models with long\-term memory\.InProceedings of the AAAI conference on artificial intelligence,Vol\.38,pp\. 19724–19731\.Cited by:[§1](https://arxiv.org/html/2609.19167#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.19167#S2.SS1.p1.1)\.
- \[49\]W\. Zhou, K\. Cao, H\. Zheng, X\. Zheng, M\. Liu, P\. O\. Kristensson, W\. W\. Mayol\-Cuevas, F\. Zhang, W\. Lin, and J\. Shen\(2025\)X\-lebench: a benchmark for extremely long egocentric video understanding\.InConference on Empirical Methods in Natural Language Processing,External Links:[Link](https://api.semanticscholar.org/CorpusID:275470689)Cited by:[§2\.2](https://arxiv.org/html/2609.19167#S2.SS2.p1.1)\.
- \[50\]Y\. Zhu, D\. Tian, and F\. Yan\(2020\)Effectiveness of entropy weight method in decision\-making\.Mathematical Problems in Engineering2020,pp\. 1–5\.External Links:[Link](https://api.semanticscholar.org/CorpusID:216219148)Cited by:[§4](https://arxiv.org/html/2609.19167#S4.p4.1)\.

\\thetitle

Supplementary Material

## Appendix AEthical Considerations

The design of ReaLMem, including its data collection, processing, and annotation, underwent ethical review and was approved by the ethical review board of the authors’ organization\. Prior to participation, each contributor was fully informed of the data\-collection procedure, the categories of personal content involved, and the intended research use, and signed an informed\-consent agreement\. Beyond consent, every record passes through the anonymization pipeline of Appendix[B](https://arxiv.org/html/2609.19167#A2)\(face blurring, PII pseudonymization, and metadata sanitization\) before any downstream use\. To guard against potential negative applications, ReaLMem will be distributed for research purposes only and released exclusively to users who sign a data license agreement governing acceptable use\.

## Appendix BDataset Anonymization

To protect participant privacy before any processing, we detect faces with RetinaFace\[[5](https://arxiv.org/html/2609.19167#bib.bib50)\]at a detection confidence threshold of0\.50\.5, and apply Gaussian blur to every detected face region\. In parallel, personally identifiable information in text—names, precise locations, and contact details—is pseudonymized with stable placeholders \(*e\.g*\., “Friend A”, “\[Restaurant B\]”\) that preserve referential consistency for analysis while preventing re\-identification\. Finally, we sanitize the raw metadata so that the released dataset contains only anonymized, privacy\-safe records\.

## Appendix CExperimental Setup Details

#### Memory system configurations\.

Mem0\[[3](https://arxiv.org/html/2609.19167#bib.bib13)\], MemOS\[[27](https://arxiv.org/html/2609.19167#bib.bib10)\], and EverMemOS\[[15](https://arxiv.org/html/2609.19167#bib.bib11)\]are run under each system’s default configuration during the memory construction stage; we replace their answering backbone with a unified GPT\-4\.1\-mini to isolate the contribution of the memory architecture from the underlying MLLM\. The resulting per\-query inference context ranges from 1\.0k \(Mem0\) to 1\.7k \(EverMemOS\) tokens\.

#### Oracle reference setup\.

The Oracle setting supplies each frontier MLLM with only the gold\-evidence sessions for the current question, identified by the per\-QA evidence annotation\. Oracle is restricted to T1 because T2 and T3 require holistic synthesis across the full participant history rather than retrieval of a localized evidence span\.

## Appendix DEvaluation Details

#### T1 Factual Memory Recall — judge prompt\.

T1 responses are evaluated by GPT\-4o\-mini asCorrectorWrong\. The judge is provided with the question, the reference answer, and the model’s response and applies four criteria: \(1\)Conceptual Match— paraphrase and non\-essential omissions are accepted as long as core entities are preserved; \(2\)Time & Dates— relative and format\-variant references are accepted if they resolve to the same calendar date \(year, month, day\); time\-of\-day differences are ignored; \(3\)Numbers & Counts— summaries requiring exact counts must match the reference; \(4\)No Contradictory Hallucinations— a response that includes correct information alongside confidently stated contradictory fabrications is markedWrong\. The T1 score for an instance is 1 \(Correct\) or 0 \(Wrong\); subtask scores are the mean over QA pairs\.

Table 7:Taxonomy of the ReaLMem benchmark\. Nine subtasks are organized into three cognitive categories spanning factual memory recall, persona inference, and predictive personalization\. QA pair counts per subtask are shown in parentheses\.
#### T2 Persona Inference — scoring rubric\.

Each T2 instance provides a question, a reference answer, and participant\-validated key supporting points\. GPT\-4o\-mini assigns two independent integer scores on a 1–5 scale:

- •Coverage\(1–5\): degree to which the response addresses the key supporting points — 5 covers all comprehensively, 3 covers roughly half, 1 misses the point entirely\.
- •Accuracy\(1–5\): logical soundness and factual consistency with the reference — 5 is fully accurate with no hallucinations, 3 has minor misinterpretations, 1 is completely fabricated or contradictory\.

The reported T2 score for an instance is\(Coverage\+Accuracy\)×10\(\\text\{Coverage\}\+\\text\{Accuracy\}\)\\times 10, giving a range of\[20,100\]\[20,100\]\.

#### T3 Predictive Personalization — ranking metric\.

Each T3 instance presents a set of candidate options and requires the model to produce a ranking\. Evaluation uses Kendall\-τ\\taurank correlation between the predicted and participant\-validated ground\-truth ranking:

τ=\(\# concordant pairs\)−\(\# discordant pairs\)\(n2\),\\tau=\\frac\{\(\\text\{\\\# concordant pairs\}\)\-\(\\text\{\\\# discordant pairs\}\)\}\{\\binom\{n\}\{2\}\},which is then affinely rescaled to\[0,100\]\[0,100\]viascore=50⋅\(τ\+1\)\\text\{score\}=50\\cdot\(\\tau\+1\), so that random ranking yields 50 and a perfect reversal yields 0\.

#### Aggregation\.

Within each subtask, scores are averaged uniformly over QA pairs\. Each cognitive level \(T1/T2/T3\) reports the QA\-count\-weighted mean over its constituent subtasks\.

## Appendix EQA Generation Details

We provide the full construction procedure for the 1,629 QA pairs inReaLMem, summarised in §3\.2 of the main text\.

#### T1 Factual Memory Recall\.

T1 QA pairs are generated from objectively annotated memory events\. Each instance is formulated as an open\-ended question accompanied by a reference answer and explicit grounding evidence \(data IDs\)\. To ensure linguistic naturalness and evaluation validity, all generated QA pairs undergo multi\-stage human quality control by trained annotators, who address unnatural phrasing, insufficient contextual constraints, improper temporal expressions, and low\-information or ambiguous answers, yielding high\-quality, well\-specified queries\.

#### T2 Persona Inference\.

T2 QA generation leverages subjective annotations capturing user preferences and personal contexts\. Because these tasks are inherently subjective, each QA pair consists of a question, a reference answer, and a set of key supporting points used for graded evaluation, rather than a single rigid ground truth\. Crucially, these instances are verified and refined by the original data contributors, who assess the completeness, correctness, and relevance of the answer against their own lived experiences, ensuring that subjective reasoning tasks remain both reliable and faithful to real user perspectives\.

#### T3 Predictive Personalization\.

For T3, we move beyond conventional multiple\-choice QA and adopt a ranking\-based evaluation paradigm that better reflects real\-world decision\-making\. Instead of collapsing preference modeling and application into binary or single\-choice outputs, each query presents multiple candidate options constructed from different combinations of user preferences, and the task requires models to produce a ranked ordering aligned with the user’s actual preference distribution\. The ground\-truth rankings are directly provided and validated by the participants, capturing the nuanced, multi\-dimensional, and context\-dependent nature of human decision\-making\. This design avoids oversimplified preference matching and instead evaluates whether models can integrate heterogeneous signals and perform flexible, personalized reasoning\.

After manual screening and revision, the pipeline yields a total of 1,629 QA pairs that are grounded, diverse, and reliable, while fully leveraging the advantages of real\-world data and human\-in\-the\-loop validation\.

## Appendix FChronoProfiler Implementation Details

#### Update operations\.

During incremental profile construction \(Algorithm[1](https://arxiv.org/html/2609.19167#alg1)\), each newly extracted candidate is compared against existing entries in the same semantic category and resolved by one of four operations:Newinserts a previously unseen attribute as a fresh entry;Duplicatekeeps the existing entry unchanged and only appends the current session to its evidence list;Updateoverwrites the entry with refined or more specific content while extending its evidence list;Conflict\(a contradiction or change relative to an existing entry,*e\.g*\.“used to like coffee” vs\. “switched to tea”\) does not overwrite the prior entry; instead it instantiates the new attribute as a separate entry that links back to the conflicting one via a related reference, so that the outdated and current signals are both retained as independently timestamped items, their competition left to be arbitrated downstream by the Temporal Stability score rather than by destructive overwriting\.

#### Hyperparameters\.

The decay constantsτs\\tau\_\{s\},τd\\tau\_\{d\}, andτr\\tau\_\{r\}control how quickly each lifecycle feature saturates\. We setτs=3\\tau\_\{s\}=3, so that support saturates around three independent mentions;τd=150\\tau\_\{d\}=150days, so duration grows substantially over roughly five months of consistent expression; andτr=60\\tau\_\{r\}=60days, giving recency a half\-life on the order of two months\. Unlike the feature timescales, the combination weights𝐰=\(ws,wd,wr\)\\mathbf\{w\}=\(w\_\{s\},w\_\{d\},w\_\{r\}\)are*not*fixed hyperparameters: they are recomputed per user by the entropy weight method over that user’s profile\-item feature matrix, so a feature that fails to discriminate among a user’s items is automatically down\-weighted\. The resulting entropy\-weighted sum is then min–max normalized within each user to\[0,1\]\[0,1\]\. The reference timeTrefT\_\{\\mathrm\{ref\}\}is set to a common evaluation timestamp shared across users \(here we set it to 2026\-01\-12\)\. At inference, profile items are retrieved per semantic category by dense query similarity \(computed with OpenAI’s text\-embedding\-3\-small\[[32](https://arxiv.org/html/2609.19167#bib.bib42)\]\) and truncated to the top\-15 per category \(depth sensitivity is analyzed in Appendix[G](https://arxiv.org/html/2609.19167#A7)\)\.

#### Profile\-swap evaluation protocol\.

For the ChronoProfiler results in Tab\. 5 and 6 of the main paper, we hold EverMemOS’s episodic retrieval and the reader backbone fixed and substitute only the injected user profile, replacing EverMemOS’s built\-in profile with our temporally\-weighted profile\. This yields a per\-query context of∼6\{\\sim\}6k tokens and forms a controlled contrast against theEverMemOS: Ep\.\+Prof\.baseline \(identical episodes and backbone, EverMemOS’s own profile\), so that any difference isolates the effect of the profile representation and its temporal stability weighting from both episodic retrieval and raw long\-context access\.

Algorithm 1Incremental Profile Construction0:Session sequence

𝒟=\{d1,…,dN\}\\mathcal\{D\}=\\\{d\_\{1\},\\dots,d\_\{N\}\\\}in chronological order\.

0:Profile

𝒫\\mathcal\{P\}with timestamped evidence\.

1:

𝒫←\\mathcal\{P\}\\leftarrowInitialize empty profile with 5 semantic categories

2:foreach session

dt∈𝒟d\_\{t\}\\in\\mathcal\{D\}do

3:

ℰt←LLM\.Extract​\(dt\)\\mathcal\{E\}\_\{t\}\\leftarrow\\text\{LLM\.Extract\}\(d\_\{t\}\)
4:foreach candidate

e∈ℰte\\in\\mathcal\{E\}\_\{t\}do

5:

relation←LLM\.Compare\(e,𝒫\[e\.category\]\)relation\\leftarrow\\text\{LLM\.Compare\}\(e,\\mathcal\{P\}\[e\.category\]\)
6:if

r​e​l​a​t​i​o​n==relation==Newthen

7:

𝒫\[e\.category\]\.Append\(e,evidence=\{dt\}\)\\mathcal\{P\}\[e\.category\]\\text\{\.Append\}\(e,\\text\{evidence\}=\\\{d\_\{t\}\\\}\)
8:elseif

r​e​l​a​t​i​o​n==relation==Duplicatethen

9:

m​a​t​c​h​\.evidence\.Add​\(dt\)match\\text\{\.evidence\.Add\}\(d\_\{t\}\)
10:elseif

r​e​l​a​t​i​o​n==relation==Updatethen

11:

m​a​t​c​h​\.Update​\(e\)match\\text\{\.Update\}\(e\);

m​a​t​c​h​\.evidence\.Add​\(dt\)match\\text\{\.evidence\.Add\}\(d\_\{t\}\)
12:elseif

r​e​l​a​t​i​o​n==relation==Conflictthen

13:

𝒫\[e\.category\]\.Append\(e,evidence=\{dt\},related\_id=match\.id\)\\mathcal\{P\}\[e\.category\]\\text\{\.Append\}\(e,\\text\{evidence\}=\\\{d\_\{t\}\\\},\\text\{related\\\_id\}=match\.id\)
14:endif

15:endfor

16:

𝒫​\.SaveCheckpoint​\(\)\\mathcal\{P\}\\text\{\.SaveCheckpoint\}\(\)
17:endfor

18:return

𝒫\\mathcal\{P\}

Algorithm 2Temporal Stability Scoring0:Profile

𝒫\\mathcal\{P\}, Reference time

TrefT\_\{\\mathrm\{ref\}\}, Decay constants

τs=3,τd=150,τr=60\\tau\_\{s\}=3,\\tau\_\{d\}=150,\\tau\_\{r\}=60\.

0:Profile

𝒫\\mathcal\{P\}with TS score for each item\.

1:foreach item

p∈𝒫p\\in\\mathcal\{P\}do

2:

𝒮p←\|Unique​\(p​\.mentioned\_times\)\|\\mathcal\{S\}\_\{p\}\\leftarrow\|\\text\{Unique\}\(p\\text\{\.mentioned\\\_times\}\)\|
3:

tfirst,tlast←MinMax​\(p​\.mentioned\_times\)t\_\{\\mathrm\{first\}\},t\_\{\\mathrm\{last\}\}\\leftarrow\\text\{MinMax\}\(p\\text\{\.mentioned\\\_times\}\)
4:

ssup←1−exp\(−𝒮p/τs\)s\_\{\\mathrm\{sup\}\}\\leftarrow 1\-\\exp\(\-\\mathcal\{S\}\_\{p\}/\\tau\_\{s\}\)
5:

sdur←1−exp\(−\(tlast−tfirst\)/τd\)s\_\{\\mathrm\{dur\}\}\\leftarrow 1\-\\exp\(\-\(t\_\{\\mathrm\{last\}\}\-t\_\{\\mathrm\{first\}\}\)/\\tau\_\{d\}\)
6:

srec←exp\(−\(Tref−tlast\)/τr\)s\_\{\\mathrm\{rec\}\}\\leftarrow\\exp\(\-\(T\_\{\\mathrm\{ref\}\}\-t\_\{\\mathrm\{last\}\}\)/\\tau\_\{r\}\)
7:endfor

8:

𝐰←EntropyWeights​\(\{\(ssup,sdur,srec\)p\}p∈𝒫\)\\mathbf\{w\}\\leftarrow\\text\{EntropyWeights\}\(\\\{\(s\_\{\\mathrm\{sup\}\},s\_\{\\mathrm\{dur\}\},s\_\{\\mathrm\{rec\}\}\)\_\{p\}\\\}\_\{p\\in\\mathcal\{P\}\}\)\{per\-user\}

9:foreach item

p∈𝒫p\\in\\mathcal\{P\}do

10:

zp←ws⋅ssup\+wd⋅sdur\+wr⋅srecz\_\{p\}\\leftarrow w\_\{s\}\\cdot s\_\{\\mathrm\{sup\}\}\+w\_\{d\}\\cdot s\_\{\\mathrm\{dur\}\}\+w\_\{r\}\\cdot s\_\{\\mathrm\{rec\}\}
11:endfor

12:

p​\.TS←\(zp−minq⁡zq\)/\(maxq⁡zq−minq⁡zq\),∀pp\\text\{\.TS\}\\leftarrow\(z\_\{p\}\-\\min\_\{q\}z\_\{q\}\)/\(\\max\_\{q\}z\_\{q\}\-\\min\_\{q\}z\_\{q\}\),\\ \\forall p\{within\-user min–max\}

13:return

𝒫\\mathcal\{P\}

Figure 3:ChronoProfiler retrieval\-depth sensitivity for the two GPT readers\. Each panel plots the temporally\-weighted profile under per\-category RAG as a function of retrieval depthkk\(solid, hollow markers\), with the full no\-RAG profile shown as the right\-most*Full*tick \(filled markers, shaded\)\. \(a\) T2 persona understanding \(LLM\-judge00–1010,×10\\times 10\); \(b\) T3 context\-aware reasoning \(Kendall’sτ\\tau,×100\\times 100\)\.

## Appendix GChronoProfiler Retrieval\-Depth Sensitivity

We fix the per\-category retrieval depth atk=15k=15and report the results in Tab\. 5 and 6 of the main paper\. Figure[3](https://arxiv.org/html/2609.19167#A6.F3)traces how our temporally\-weighted profile behaves under per\-category RAG askkvaries, for the two GPT readers, with the full no\-RAG profile \(∼16\{\\sim\}16k tokens\) shown as the right\-most categorical tick for reference\.

#### T2 saturates well before the full budget\.

Persona understanding \(T2\) rises smoothly withkkand is essentially flat byk=15k=15, where both readers sit within∼0\.8\{\\sim\}0\.8points of their full\-profile score \(GPT\-4\.1\-mini:78\.378\.3vs\.78\.578\.5; GPT\-5\.4\-mini:79\.879\.8vs\.79\.079\.0\) while using only∼6\{\\sim\}6k of the∼16\{\\sim\}16k full\-profile tokens \(≈40%\{\\approx\}40\\%\)\. Retrieving more profile items therefore buys little on persona inference, which motivates our top\-15 default\.

#### T3 is reader\-dependent\.

Context\-aware prediction \(T3\) does not share a single optimum\. GPT\-4\.1\-mini peaks at the RAG operating point \(k=15k=15,61\.461\.4\) and*declines*when given the full profile \(59\.359\.3\), consistent with the dialogue\-dilution effect we observe in the modality ablation \(§5\.3\): a longer, noisier preference record introduces conflicting signals that this reader cannot arbitrate\. GPT\-5\.4\-mini, in contrast, continues to benefit from the full profile \(62\.562\.5vs\.60\.860\.8atk=15k=15\)\. Because the compact top\-15 profile is at or near the best T3 setting for GPT\-4\.1\-mini and within∼1\.7\{\\sim\}1\.7points for GPT\-5\.4\-mini, at a fraction of the token cost, we adoptk=15k=15as the single default across readers\.

Similar Articles

MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale

arXiv cs.CL

MemArena is a new ego-centric benchmark for evaluating on-device personal memory assistants, using a MASim agent simulator to generate multi-session conversational worlds and ground truth across recall, reasoning, and trustworthiness dimensions. Initial results show memory-backend choice often matters more than reader scale, and permission-aware access remains a universal challenge.

MemoryForge: Synthesize Lifelong Memory for Human-Like LLM Agents

arXiv cs.CL

This paper introduces MemoryForge, a framework for synthesizing lifelong autobiographical memory from brief target personas to enable frozen LLMs to exhibit more human-like behaviors in role-play and user-simulation, outperforming descriptive conditioning baselines.