DREAM: LLM-based Dynamic Role-playing via Event-Aware Memory Graph

arXiv cs.CL Papers

Summary

This paper introduces DREAM, a structured memory framework for LLM-based role-playing agents that uses an Event-aware Memory Graph to maintain temporal and causal coherence, and proposes the TCM benchmark for evaluation.

arXiv:2608.05170v1 Announce Type: new Abstract: Role-playing agents (RPAs) have emerged as a key application of large language models, enabling immersive and high-fidelity character simulation. Accurate role-playing of established characters requires not only stylistic imitation but also temporally consistent and causally grounded behavioral reasoning. However, existing RPAs primarily rely on static character descriptions and unstructured memory, limiting their ability to maintain long-term narrative and personality coherence. We introduce DREAM, a structured memory framework for role-playing agents inspired by the Activating Event-Belief-Consequence (ABC) cognitive model. DREAM transforms unstructured literary text into an Event-aware Memory Graph (EMG) that organizes character experiences into temporally ordered and causally linked event graph. This representation enables the construction of dynamic, dual-granularity character profiles that capture both stable personality traits and event-driven behavioral evolution. We further propose the Temporal Causal Memory (TCM) benchmark to evaluate temporal consistency and long-range causal narrative coherence. DREAM achieves state-of-the-art performance across CoSER, LIFECHOICE, and TCM, outperforming multiple strong baselines. Our approach demonstrates the effectiveness of structured memory in enhancing the interpretability and consistency of role-playing agents.
Original Article
View Cached Full Text

Cached at: 08/07/26, 07:50 AM

# DREAM: LLM-based Dynamic Role-playing via Event-Aware Memory Graph
Source: [https://arxiv.org/html/2608.05170](https://arxiv.org/html/2608.05170)
Zhihao Xiao[0009\-0009\-5779\-1345](https://orcid.org/0009-0009-5779-1345)Hangzhou International Innovation Institute, Beihang UniversityHangzhouChina[xiaozhihao@buaa\.edu\.cn](https://arxiv.org/html/2608.05170v1/mailto:[email protected])Mengting LiHangzhou International Innovation Institute, Beihang UniversityHangzhouChina[li˙mengting@buaa\.edu\.cn](https://arxiv.org/html/2608.05170v1/mailto:li%CB%[email protected]),Xintao WangSchool of Computer Science, Fudan UniversityShanghaiChina[xtwang21@m\.fudan\.edu\.cn](https://arxiv.org/html/2608.05170v1/mailto:[email protected]),Linfeng LiHangzhou International Innovation Institute, Beihang UniversityHangzhouChina[zy2457211@buaa\.edu\.cn](https://arxiv.org/html/2608.05170v1/mailto:[email protected]),Limin ShuiHangzhou International Innovation Institute, Beihang UniversityHangzhouChina[Liminshui@buaa\.edu\.cn](https://arxiv.org/html/2608.05170v1/mailto:[email protected]),Mengqi JiHangzhou International Innovation Institute, Beihang UniversityHangzhouChina[jimengqi@buaa\.edu\.cn](https://arxiv.org/html/2608.05170v1/mailto:[email protected])andBorui CaiHangzhou International Innovation Institute, Beihang UniversityHangzhouChina[caibr@buaa\.edu\.cn](https://arxiv.org/html/2608.05170v1/mailto:[email protected])

\(2026\)

###### Abstract\.

Role\-playing agents \(RPAs\) have emerged as a key application of large language models, enabling immersive and high\-fidelity character simulation\. Accurate role\-playing of established characters requires not only stylistic imitation but also temporally consistent and causally grounded behavioral reasoning\. However, existing RPAs primarily rely on static character descriptions and unstructured memory, limiting their ability to maintain long\-term narrative and personality coherence\. We introduce DREAM, a structured memory framework for role\-playing agents inspired by the Activating Event–Belief–Consequence \(ABC\) cognitive model\. DREAM transforms unstructured literary text into an Event\-aware Memory Graph \(EMG\) that organizes character experiences into temporally ordered and causally linked event graph\. This representation enables the construction of dynamic, dual\-granularity character profiles that capture both stable personality traits and event\-driven behavioral evolution\. We further propose the Temporal Causal Memory \(TCM\) benchmark to evaluate temporal consistency and long\-range causal narrative coherence\. DREAM achieves state\-of\-the\-art performance across CoSER, LIFECHOICE, and TCM, outperforming multiple strong baselines\. Our approach demonstrates the effectiveness of structured memory in enhancing the interpretability and consistency of role\-playing agents\.

Knowledge Graph, Role\-playing Agents, Large Language Models, Causal Narrative Modeling, Temporal Memory Retrieval

††copyright:none††journalyear:2026††copyright:cc††conference:Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V\.2; August 09–13, 2026; Jeju Island, Republic of Korea††booktitle:Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V\.2 \(KDD ’26\), August 09–13, 2026, Jeju Island, Republic of Korea††doi:10\.1145/3770855\.3818027††isbn:979\-8\-4007\-2259\-2/2026/08††ccs:Computing methodologies Natural language processing††ccs:Computing methodologies Knowledge representation and reasoning## 1\.Introduction

Recent advances in large language models \(LLMs\) have enabled a new generation of role\-playing agents \(RPAs\) that can simulate the linguistic styles and reasoning patterns of specific characters\(Tsenget al\.,[2024](https://arxiv.org/html/2608.05170#bib.bib23); Chenet al\.,[2024a](https://arxiv.org/html/2608.05170#bib.bib24)\)\. Such agents have shown promise in immersive games\(Xuet al\.,[2025a](https://arxiv.org/html/2608.05170#bib.bib37)\), interactive fiction, and virtual companions, where maintaining role consistency over long and dynamic interactions is critical\(Parket al\.,[2023](https://arxiv.org/html/2608.05170#bib.bib25); Zhouet al\.,[2023](https://arxiv.org/html/2608.05170#bib.bib26)\)\. However, despite recent progress\(Zhouet al\.,[2024](https://arxiv.org/html/2608.05170#bib.bib1); Wanget al\.,[2024](https://arxiv.org/html/2608.05170#bib.bib7)\), constructing RPAs for established characters remains challenging\. Beyond surface\-level persona imitation, effective role\-playing requires agents to align their responses with a character’s past experiences, cognition, and motivations as shaped by the underlying narrative over time, as shown in Fig\.[1](https://arxiv.org/html/2608.05170#S1.F1)\.

RPAs typically rely on two kinds of persona data: profiles and memories\. However, existing approaches exhibit two key limitations in effectively leveraging resources: \(1\)Static Profiles:Most existing methods inject predefined character profiles designed for fixed scenarios through prompting or fine\-tuning\(Luet al\.,[2024](https://arxiv.org/html/2608.05170#bib.bib3); Zhouet al\.,[2024](https://arxiv.org/html/2608.05170#bib.bib1); Heet al\.,[2025](https://arxiv.org/html/2608.05170#bib.bib2)\)\. While effective at capturing surface\-level stylistic traits, such representations fail to reflect how a character’s disposition is shaped and updated by past experiences\. As a result, RPAs often respond rigidly and struggle to adapt to dynamic interactions, where the profile of a character should be dynamically grounded in its memory rather than predefined scenarios\. \(2\)Fragmented Memory:Prior work on character memory typically follows two paradigms: retrieving past dialogues or events based on semantic similarity\(Liet al\.,[2023](https://arxiv.org/html/2608.05170#bib.bib6); Xuet al\.,[2025b](https://arxiv.org/html/2608.05170#bib.bib8)\), or encoding experiences implicitly within model parameters\(Wanget al\.,[2025](https://arxiv.org/html/2608.05170#bib.bib10)\)\. Although these approaches enable memory access, they organize character memory as isolated units, thereby fragmenting complete plots, leaving temporal order and inter\-event causal relationships implicit\. Consequently, RPAs can recall what happened, but cannot reliably reason about when and why past experiences should influence current behavior, leading to narrative inconsistency and temporal leakage\(Ahnet al\.,[2024](https://arxiv.org/html/2608.05170#bib.bib27)\)\. Moreover, parametric memories often discard fine\-grained event details and offer limited interpretability, further constraining their applicability in character\-centric reasoning\.

To address these challenges, we propose DREAM \(DynamicRole\-playing viaEvent\-AwareMemory\), a role\-playing multi\-agent system with event\-aware memory and dynamic profiles\. DREAM explicitly organizes character experiences as an Event\-aware Memory Graph \(EMG\), where events are organized by temporal order and causal dependencies\. In EMG, each event encodes both global narrative context and fine\-grained character dynamics, enabling agents to reason over how past experiences shape character development\. During interaction, DREAM retrieves temporally valid memories from the EMG to dynamically synthesize event\-based character profiles, ensuring response generation remains consistent with the character’s time\-dependent cognitive state in the storyline\.

Specifically, DREAM constructs the EMG by extracting information from literary texts at dual granularities, i\.e\., global narrative context at the macro level and fine\-grained dynamics at the micro level\. At the macro level, it captures event background, world settings, and character attributes\. At the micro level, it adopts causal narrative chains to capture fine\-grained emotion\-cognition\-behavior dynamics within unit plots, inspired by the well\-established psychological Activating Event\-Belief\-Consequence \(ABC\) model\(Ellis,[1957](https://arxiv.org/html/2608.05170#bib.bib36)\)\. During role\-playing, DREAM employs a temporally constrained hybrid retrieval that integrates semantic matching with graph\-based multi\-hop inference; this mechanism enables the agent to trace the causal origins of a character’s behavior and evolution, ensuring that responses are grounded in a logically coherent and chronologically valid narrative context\.

To complement evaluation dimensions that are underexplored in prior work, we propose the TCM benchmark, which focuses on two key aspects: \(1\) coherent temporal memory and \(2\) long\-term causal narrative consistency\. We conduct comprehensive experiments on CoSER, LIFECHOICE, and TCM, and the results show that DREAM achieves state\-of\-the\-art performance across baselines\.

Our contributions are summarized as follows:

- •We introduce DREAM, a multi\-agent framework for role\-playing that systematically constructs memory graphs and generates dynamic character profiles to support context\-aware interactions\.
- •At the core of DREAM, we propose the Event\-aware Memory Graph \(EMG\), which encodes temporally ordered events and causal dependencies, enabling dynamic profile construction and contextually grounded retrieval\.
- •Extensive experiments on various benchmarks demonstrate that DREAM achieves state\-of\-the\-art performance across baselines, consistently improving temporal narrative consistency and role\-playing ability\.

![Refer to caption](https://arxiv.org/html/2608.05170v1/x1.png)Figure 1\.A role\-playing example of the DREAM framework\. The framework leverages the EMG and temporal filtering to construct dynamic profiles, enabling LLMs to generate contextually accurate and emotionally resonant interactions, such as Harry Potter’s encounter with the Mirror of Erised\.
## 2\.Related Work

### 2\.1\.Role\-Playing Agents with LLMs

Role\-playing agents \(RPAs\) aim to achieve high\-fidelity simulations of specific personas\. Prior research has mainly focused on: \(1\)Profile Construction:Injecting static profiles via SFT or prompting\.Ditto\(Luet al\.,[2024](https://arxiv.org/html/2608.05170#bib.bib3)\)utilizes self\-alignment mechanisms to elicit intrinsic character knowledge from LLMs, whereasCharacterGLM\(Zhouet al\.,[2024](https://arxiv.org/html/2608.05170#bib.bib1)\)integrates diverse sources including human role\-playing and literature extraction\. To further enhance persona depth, profiles are often enriched with psychological attributes\(Occhipintiet al\.,[2024](https://arxiv.org/html/2608.05170#bib.bib4)\)or multimodal information\(Daiet al\.,[2024](https://arxiv.org/html/2608.05170#bib.bib5); Zhanget al\.,[2025a](https://arxiv.org/html/2608.05170#bib.bib21)\)\. Moreover,Crab\(Heet al\.,[2025](https://arxiv.org/html/2608.05170#bib.bib2)\)introduces a configurable framework to improve profile flexibility\. \(2\)Cognitive Alignment:Constraining the model’s internal reasoning to align its cognitive mechanisms with the target character\. Tang et al\.\(Tanget al\.,[2025](https://arxiv.org/html/2608.05170#bib.bib11)\)employ character\-centric Chain\-of\-Thought \(CoT\) to explicitly guide reasoning paths\. Alternatively, reinforcement learning approaches internalize motivations through contrastive preferences\(Yeet al\.,[2025](https://arxiv.org/html/2608.05170#bib.bib12)\), dual\-process architectures\(Liuet al\.,[2025](https://arxiv.org/html/2608.05170#bib.bib13)\), or multi\-objective alignment\(Liaoet al\.,[2025](https://arxiv.org/html/2608.05170#bib.bib14)\)\. However, existing studies predominantly rely on static character profiles, limiting their ability to adapt behaviors across evolving interaction contexts\. This limitation motivates research on incorporating event\-aware memory mechanisms into RPAs\.

### 2\.2\.Agent Memory Systems

Memory systems for LLM\-based agents\(Zhanget al\.,[2025b](https://arxiv.org/html/2608.05170#bib.bib17); Huet al\.,[2025](https://arxiv.org/html/2608.05170#bib.bib15)\)primarily rely on parametric storage within model weights\(Fanget al\.,[2024](https://arxiv.org/html/2608.05170#bib.bib18)\)or external retrieval paradigms\(Yanet al\.,[2025](https://arxiv.org/html/2608.05170#bib.bib19)\)\. Notably, GraphRAG\(Edgeet al\.,[2024](https://arxiv.org/html/2608.05170#bib.bib16)\)advances the latter by leveraging graph\-based indexing to capture both local and global semantic relationships\. Recently, specialized memory mechanisms for RPAs have emerged\(Chenet al\.,[2023](https://arxiv.org/html/2608.05170#bib.bib9); Wanget al\.,[2024](https://arxiv.org/html/2608.05170#bib.bib7)\)\. For instance, ChatHaruhi\(Liet al\.,[2023](https://arxiv.org/html/2608.05170#bib.bib6)\)retrieves relevant past dialogs via similarity\-based retrieval, which primarily captures conversational style\. CHARMAP\(Xuet al\.,[2025b](https://arxiv.org/html/2608.05170#bib.bib8)\)stores key events in a key\-value memory bank, yet its chunk\-based extraction may fragment complete plots\. CoSER\(Wanget al\.,[2025](https://arxiv.org/html/2608.05170#bib.bib10)\)further incorporates character motivations at the event level, but it overlooks the causal dependencies between those events\. Overall, these methods suffer from fragmented memory organization, limited interpretability, and a lack of explicit inter\-event causal structure, resulting in largely static character representations\. In contrast, we construct explicit character memories at the event granularity, organizing experiences into structured causal chains to support interpretable and consistent long\-term role\-playing\.

## 3\.The DREAM Framework

In this section, we elaborate on the design of DREAM\. The primary goal of DREAM is to support dynamic role\-playing through the EMG, with two key objectives:

- •Establish the EMG as a structured memory that encodes the character’s experiences into causally linked event chains, facilitating a holistic understanding of their behavioral logic\.
- •Instantiate dynamic character profiles through EMG\-based retrieval, effectively mapping the character’s time\-dependent cognitive state throughout the interaction\.

The design of DREAM stems from a requirement that high\-quality RPAs must possess memory continuity and narrative depth\. DREAM follows a two\-stage pipeline: first, it transforms unstructured literary texts into the structured EMG that encodes temporal and causal dependencies\. During role\-playing, it performs targeted retrieval from this EMG to instantiate dynamic, context\-aware character profiles for response generation\.

### 3\.1\.Core Formulation of DREAM

DREAM enables dynamic and experience\-driven role\-playing by maintaining character memory with explicit temporal and causal structure\. This capability is supported by the Event\-Aware Memory Graph \(EMG\), which organizes fragmented character experiences into a structured memory graph with timestamps and causal dependencies\.

#### 3\.1\.1\.Problem Formulation\.

In a literary role\-playing setting, given a literary corpusℬ\\mathcal\{B\}and a target charactercc, our aim is to optimize LLMs to simulate the behavior and personalityPPofcc\. We define a character as a dual\-structured representation:

\(1\)𝒫c=\(​𝒜​,𝒢e​v​e​n​t​\)\\mathcal\{P\}\_\{c\}=\(\\text\{ \}\\mathcal\{A\}\\text\{ \},\\mathcal\{G\}\_\{event\}\\text\{ \}\)where𝒜\\mathcal\{A\}denotes the identity attributes of the character \(e\.g\., background, social roles, and inherent personality traits\), and𝒢e​v​e​n​t\\mathcal\{G\}\_\{event\}represents an EMG\. Unlike traditional flat memory banks,𝒢e​v​e​n​t\\mathcal\{G\}\_\{event\}is a structured representation extracted from original text that encodes character persona, experiences, causal dependencies, and psychological evolution\.

Grounding interactions in𝒢e​v​e​n​t\\mathcal\{G\}\_\{event\}, the DREAM framework follows a retrieve\-generate\-respond workflow\. At each dialogue turn, we utilize the user queryQQas a probe to traverse𝒢e​v​e​n​t\\mathcal\{G\}\_\{event\}and retrieve: \(1\) anchor event, the event most relevant toQQand the scenarioSS; and \(2\) causal chains, which represent the causal narrative chains encapsulating character cognition and motivations within anchor events\. DREAM performs a two\-stage process to generate a dynamic profile:

- •Event Information Retrieval \(ℛe​v​e​n​t\\mathcal\{R\}\_\{event\}\): Retrieve the event\-based character personalities that most fit the current scenario by traversing the timestampstit\_\{i\}\.
- •Dynamic Profile Instantiation \(𝒫d​y​n\\mathcal\{P\}\_\{dyn\}\): Based on the retrieved context, we model the character’s time\-dependent cognitive state, belief updates, and personality evolution\.

The process is formalized as follows:

\(2\)𝒫d​y​n​a​m​i​c=Generate​\(Re​v​e​n​t​\(Q,S\)\)\\mathcal\{P\}\_\{dynamic\}=\\text\{Generate\}\(\\text\{R\}\_\{event\}\(Q,S\)\)Before each response is generated by the LLM \(Eq\. \([3](https://arxiv.org/html/2608.05170#S3.E3)\)\), DREAM adopts the above generated dynamic profile to appropriately respond to user query:

\(3\)R​e​s​p​o​n​d=L​L​M​\(Q,ℛe​v​e​n​t,𝒫d​y​n​a​m​i​c\)Respond=LLM\(Q,\\mathcal\{R\}\_\{event\},\\mathcal\{P\}\_\{dynamic\}\)

#### 3\.1\.2\.The EMG Definition

The design of EMG aims to overcome the limitations of static profiles and fragmented memory of existing role\-playing agents \(RPAs\):

EMG organizes memories with eventEiE\_\{i\}as the basic unit and connects experiences at different stages through explicit temporal and causal structures\. Then, retrieved context from the EMG can form dynamic profiles that accurately reflect the cognitive states, motivation updates, and personality evolution, thereby moving beyond surface\-level stylistic imitation\. To achieve that, we define EMG as a multi\-attribute\-directed knowledge graph:

\(4\)𝒢e​v​e​n​t=\(Em​a​c​r​o,m​i​c​r​o,Rt​e​m​p​o​r​a​l,c​a​u​s​a​l\)\\mathcal\{G\}\_\{event\}=\(E\_\{macro,micro\},R\_\{temporal,causal\}\)
Events are central entities in EMG, and we also design peripheral entities to represent detailed information of entities at macro and micro levelsEm​a​c​r​o,m​i​c​r​oE\_\{macro,micro\}\. The macro\-level includes entities that reflect event background information \(Worldview, Environments\) and character’s current status \(Identity, Social Relationships, Appearance, Habitual Actions, and Representative Dialog\)\. The micro\-level includes entities that depict fine\-grained narrative and psychological units \(Unit\-Plots, Emotions, Cognition, Behaviors, Specific Scenes, Items, and Skills\)\.

We organize relations in EMG into two typesRt​e​m​p​o​r​a​l,c​a​u​s​a​lR\_\{temporal,causal\}: temporal relations and causal relations\. Temporal relations connect events along the narrative timeline \(e\.g\.,n​e​x​t​\_​e​v​e​n​t,n​e​x​t​\_​p​l​o​tnext\\\_event,next\\\_plot\), ensuring coherent chronological memory organization\. Causal relations capture dependencies among character cognition, emotions, behaviors, and event factors \(e\.g\.,m​o​t​i​v​a​t​e​d​\_​b​ymotivated\\\_by,e​m​o​t​i​o​n​\_​f​r​o​memotion\\\_from\), modeling transitions across internal states and external actions\. In addition, EMG includes a small set of conventional relations to represent standard entity relationships\. Formal definitions of entities and relations are available in the Appendix[A](https://arxiv.org/html/2608.05170#A1)\.

### 3\.2\.EMG Construction

We develop a Memory Construction Agent to construct EMG from unstructured literary texts, such as novels and plays\. The aim is to explicitly encode temporally ordered causal events and rich character information, which are later used to derive event\-based dynamic character profiles\.

#### 3\.2\.1\.Event and Temporal Structure Construction\.

Event Construction is performed through chapter\-level narrative segmentation followed by character\-centric event categorization\. We first segment literary texts at the chapter level and employ LLMs to identify chapter boundaries and remove non\-narrative sections, which is adopted to preserve event completeness and mitigate plot fragmentation introduced by conventional fixed\-length or sentence\-level segmentation methods\.

Each chapter is then categorized based on character relevance and event completeness into complete events, incomplete events, or contextual content\.Complete event:An event including the cause, process, and outcome of the character\.Incomplete event:Chapter related to the character but insufficient to form a complete event\. In this case, next chapter is appended until a complete event can be identified\.Contextual content:Content that does not involve any character events and can serve as supplementary information, and is incorporated into the EMG\.

To enable temporally aware retrieval and prevent future leakage \(i\.e\., spoilers\), we add atit\_\{i\}to everyEiE\_\{i\}andUpU\_\{p\}node based on the extracted narrative sequence\. This ensures that RPAs can only access memories preceding the current narrative timetit\_\{i\}during role\-playing\.

![Refer to caption](https://arxiv.org/html/2608.05170v1/x2.png)Figure 2\.The structure and definition of knowledge graph![Refer to caption](https://arxiv.org/html/2608.05170v1/x3.png)Figure 3\.The complete pipeline of DREAM\. Left: DREAM is sourced from renowned books and processed via an LLM\-based pipeline\. By classifying the events of the selected character and performing dual\-granularity structured extraction, entities are obtained\. Middle: Macro\-level \(yellow\) and micro\-level \(green\) nodes are integrated to construct the EMG containing narrative chains\. Right: Timestamp filtering is used to dynamically generate character profiles, which are then used in dialogue simulations and evaluated from multiple dimensions by LLMs and humans\.
#### 3\.2\.2\.Event Internal Semantics Construction\.

Existing extraction approaches, such as dialogue\-centric extraction\(Liet al\.,[2023](https://arxiv.org/html/2608.05170#bib.bib6)\)and chunk\-level extraction\(Xuet al\.,[2025b](https://arxiv.org/html/2608.05170#bib.bib8)\), fail to jointly model global event structure and fine\-grained character dynamics over time\. To construct a structured memory framework capable of simultaneously capturing personas and the evolution ofCoC\_\{o\}and experiences in literary works, we propose dual\-granularity knowledge extraction, as shown in Fig\.[2](https://arxiv.org/html/2608.05170#S3.F2):

##### Macro\-Level Event Context

This level operates at an event\-level abstraction, capturing world settings, background information, and character attributes for theii\-th event, denoted asEiE\_\{i\}\. This structured context enables agents to maintain narrative consistency when reasoning about specific events\. We further incorporate a timestamptit\_\{i\}into the event sequence to prevent future\-event leakage during memory retrieval\. For eachEiE\_\{i\}, we define entities and relations \(partially displayed in Fig\.[3](https://arxiv.org/html/2608.05170#S3.F3), details in Table[6](https://arxiv.org/html/2608.05170#A5.T6)\) to extract a structured tuple of contextual attributes\.

##### Micro\-Level Dynamic Narrative

To capture the causal logic behind character development, we introduce fine\-grainedUpU\_\{p\}extraction\. Specifically, we defineUpU\_\{p\}as a fine\-grained sub\-unit withinEiE\_\{i\}\. Rather than simply recording what happened, we model why it happened through chains of emotion, cognition, and behavior in each associatedUpU\_\{p\}\. The extraction captures character development, emotional and cognitive changes, and behavioral causal relationships within each event\. For each event, we define entities and relations to extract a set of structured attributes and define the narrative flow as a directed graph ofUpU\_\{p\}\.

For coreference resolution during graph construction \(such as the sameEmotioninEiE\_\{i\}\), we ensure consistent entity naming via prompt\-guided extraction\. \(Detailed entity and relation definitions, along with extraction prompts, are provided in Appendix[A](https://arxiv.org/html/2608.05170#A1)\.\)

To ensure a coherent narrative structure, we employ an entity resolution strategy that automatically consolidates recurring entities across various event segments\. This maintains the continuity of character identities and integrates disparate event nodes into a unified, timeline\-based knowledge graph\.

#### 3\.2\.3\.Event Causal Structure Construction\.

To ensure consistent character memory representation and efficient retrieval, we propose a construction pipeline\. Following chronological order with roles and events as core nodes, we combine the extracted nodes into triples through predefined relationships on an event\-by\-event basis and add them to the graph until the complete EMG is built\. We employ a synchronized data ingestion strategy to efficiently fuse structured information from both macro\- and micro\-levels into the EMG \(Fig\.[3](https://arxiv.org/html/2608.05170#S3.F3)\)\.

In complex stories, characters do not act randomly; their behaviors are motivated by perceptions and experiences\. To integrate this cycle of experience, cognition, and action into the EMG, we design a causal narrative construction mechanism:

Cognitive Chain:This chain models the dynamic evolution of a character’s mental state throughout the narrative\. It is achieved by defining entities \(e\.g\.,cognition,behavior\) and relations \(e\.g\.,motivated\-by,cognition\-update\-to\) to accurately capture how cognitive states drive behavioral changes\. The red lines of the EMG in Fig\.[3](https://arxiv.org/html/2608.05170#S3.F3)show part of the Cognition Chains\.

Causal Chain:The causal chain reveals character development\. By selecting nodes \(e\.g\.,Unit\-Plot,Behavior\) and relations \(e\.g\.,next\-plot,motivated\-by\), we construct a causal chain based on the character’s motivation and experience\. This chain explicitly represents the inherent causal dependencies between character behaviors and events, providing a robust structured foundation for scenarios such as plot and behavior generation\. The green and blue lines of the EMG in Fig\.[3](https://arxiv.org/html/2608.05170#S3.F3)form a causal narrative chain\.

### 3\.3\.Memory Retrieval and Dynamic Profile Generation

This section introduces how DREAM utilizes a temporal memory retrieval mechanism to dynamically generate profiles for role\-playing\. Based on the constructed EMG, we first retrieve the subgraphs corresponding to 8\-dimensional \(shown in Fig\.[3](https://arxiv.org/html/2608.05170#S3.F3)\) personas and generate descriptions of personas via the subgraphs\.

Through given scenarios or contexts during the role\-playing, we retrieve matched event nodes and their corresponding timestamp\. Before RPAs generate responses, DREAM rewrites connected causal narrative subgraphs into descriptive paragraphs, so as to realize the dynamic update of profile during the role\-playing process\.

#### 3\.3\.1\.Hybrid Memory Retrieval Mechanism

To ensure semantic consistency, narrative logical rigor, and retrieval efficiency, while mitigating hallucinations of LLMs during long\-range role\-playing, we propose a hybrid retrieval architecture:

##### Semantic Similarity Retrieval

To retrieve relevant event information, we utilize cosine similarity to perform semantic matching between the descriptions ofEventand the current context\. This process enables DREAM to efficiently locate top\-KKmost relevant events\. And we obtain the biggesttit\_\{i\}in these events for temporal retrieval\.

##### Time\-Constrained Memory Retrieval

To retrieve Character trait information and uphold strict narrative logic with preventing future leakage, we implement a temporal memory retrieval mechanism\. This mechanism restricts the searchable memory space to events occurring before the current timestamp\. The mechanism is applied to retrieve events and specific nodes—such asUnit\-Plot, andCognition—ensuring all retrieved details are chronologically valid for the character’s current state\.

##### Structure\-based Multi\-hop Retrieval

To capture character’s long\-range causal dependencies or the abstract evolution of cognition\. We leverage the graph structure of the EMG for deep memory retrieval:

Cross\-dimensional Retrieval:By traversing relations such asmotivated\-byandcognition\-from, the agent can trace the underlying psychological drivers and cognitive shifts behind a character’s behavior\.

Causal Chain Retrieval:When query involves a character’s long developmental history or complex relationships, we implement a multi\-hop retrieval strategy, which process initiates fromUnit\-Plotand propagates along edges to harvest connectedCognitionandBehaviornodes\. This allows agents to synthesize details into a coherent causal chain \(e\.g\., linking a past psychological trauma to a current reaction\), resulting in dialogue responses with greater personality depth\.

Table 1\.The performance\(%\) of DREAM and baselines on the comprehensive role\-playing benchmark CoSER\.Underlinedvalues indicate best performance across each model,Boldvalues indicate best performance in all models\.ModelsMethodsStorylineConsistencyAnthropomorphismCharacterFidelityStorylineQualityAverageClosed\-Source LLMsGPT\-4oRAG56\.3342\.9642\.5467\.1652\.50GraphRAG59\.3445\.6243\.7171\.1654\.96GCA61\.8349\.6647\.7578\.1659\.35DREAM67\.2549\.3650\.4584\.3862\.86Gemini\-2\.5\-ProRAG58\.7043\.7945\.7469\.2654\.37GraphRAG56\.5446\.2246\.1768\.8354\.44GCA59\.8545\.8346\.9578\.8657\.87DREAM63\.2543\.3649\.7581\.3859\.44Open\-Source LLMsLLaMA\-3\.1\-8BRAG49\.7740\.7535\.9459\.7646\.56GraphRAG50\.3441\.6235\.3761\.2947\.16GCA53\.6345\.5640\.7572\.8653\.20DREAM58\.2545\.8545\.9578\.8357\.22Qwen\-2\.5\-72BRAG53\.7745\.7345\.5468\.7653\.45GraphRAG55\.3445\.4245\.1270\.2954\.04GCA58\.7248\.0851\.4974\.1058\.10DREAM60\.3649\.9649\.3579\.3859\.76Role\-Playing LLMsCharacterGLM\-6BRAG48\.3641\.6836\.7063\.8647\.65GraphRAG50\.7443\.5839\.3466\.2149\.97GCA53\.3345\.6641\.5069\.1452\.41DREAM54\.9545\.4642\.4572\.2853\.79CoSER\-8BRAG51\.6444\.6839\.7870\.5051\.65GraphRAG55\.9445\.9642\.9871\.2954\.04GCA58\.7348\.5646\.5573\.4556\.82DREAM60\.4547\.3646\.1578\.2858\.06

#### 3\.3\.2\.Dynamic Profile Generation

Considering the speed and accuracy of profile generation during dynamic role\-playing, we propose the dynamic profile constructed by rewriting the EMG information based on dynamic retrieval\.

Event\-based Information: We provide an automatic generation method based on incremental updating inspired by Hollmwood\(Chenet al\.,[2024b](https://arxiv.org/html/2608.05170#bib.bib22)\)\. We employ LLMs to automatically synthesize multi\-dimensional character descriptions derived from the EMG to construct event\-based character personalities, including background, appearance, linguistic style, key experience, core motivation, relationship, cognition chains and personality traits\.

Dynamic Profile: For the context or a given scenario in the current role\-playing, DREAM uses semantic similarity retrieval to obtain the anchor events withtit\_\{i\}and retrieves the memory subgraph through hybrid retrieval \(see[3\.3\.1](https://arxiv.org/html/2608.05170#S3.SS3.SSS1)\), and rewrites the parts of theℛe​v​e​n​t\\mathcal\{R\}\_\{event\}that need to be updated via LLMs\. Finally, the profile that is most suitable for the current context is dynamically generated\.

Table 2\.Results of 3 LLMs with 3 methods on LIFECHOICE\. ACC refers to the decision accuracy\. \+motivation refers to the results with character motivations\.MethodModelACC\+motivationProfile & MemoryGraphRAGGPT\-4o68\.1594\.25LLaMA\-3\.1\-8B61\.2292\.43CoSER\-8B69\.9495\.77CHARMAPGPT\-4o69\.3596\.92LLaMA\-3\.1\-8B65\.9295\.30CoSER\-8B70\.1495\.97DREAMGPT\-4o69\.9797\.53LLaMA\-3\.1\-8B65\.2995\.18CoSER\-8B71\.2596\.03

## 4\.Experiments

In this section, we comprehensively evaluate DREAM for role\-playing ability using the LLM\-as\-a\-judge paradigm\(Zhenget al\.,[2023](https://arxiv.org/html/2608.05170#bib.bib34); Liet al\.,[2025](https://arxiv.org/html/2608.05170#bib.bib35)\), human evaluation, and objective multiple\-choice tasks\.

### 4\.1\.Evaluation Metrics

Following CoSER\(Wanget al\.,[2025](https://arxiv.org/html/2608.05170#bib.bib10)\), we evaluate simulated role\-play conversations using GPT\-4o as a critic across four key dimensions: \(1\)Storyline Consistency, \(2\)Anthropomorphism, \(3\)Character Fidelity, and \(4\)Storyline Quality\. Assesses the naturalness of simulated conversations, focusing on narrative flow and logical consistency\. Detailed rubrics are provided in Appendix[C](https://arxiv.org/html/2608.05170#A3)\.

For the LIFECHOICE benchmark, we report decision\-making accuracy in RPAs on multiple\-choice behavioral selections and focus on two key aspects: \(1\) decision accuracy based on given context and problem \(2\) decision accuracy with given motivations\.

To supplement the indicators lacking in the existing evaluations, we propose TCM benchmark and focus on two key aspects: \(1\) coherent temporal memory, and \(2\) long\-term causal narrative consistency\.

We define two metrics to further assess simulated role\-playing dialogues: \(1\)Future Knowledge Leakage:Assesses whether dialogues avoid future leakage with respect to the current timeline\. \(2\)Causal Consistency:Assesses whether RPAs behaviors follow causal logic and are grounded in the character’s cognition and experiences\.

Table 3\.The win rate \(%\) from the TCM evaluation, Comparing DREAM Based on Memory Retrieval with RAG, GraphRAG and GCA simulation method from CoSER, where DM\. refers to DREAM, FKL\. refers to Future Knowledge Leakage and CC\. refers to Causal Consistency\.ModelMethodFKL\.CC\.Closed\-Source LLMsGPT\-4oDM\. vs RAG95\.3779\.30vs GraphRAG91\.4372\.59vs GCA82\.1958\.12Gemini\-2\.5\-ProDM\. vs RAG93\.5185\.26vs GraphRAG90\.4578\.91vs GCA85\.1963\.32Open\-Source LLMsLLaMA\-3\.1\-8BDM\. vs RAG78\.4578\.67vs GraphRAG80\.7775\.83vs GCA61\.5567\.35Qwen\-2\.5\-72BDM\. vs RAG82\.5475\.85vs GraphRAG88\.3269\.38vs GCA78\.8659\.31Role\-Playing LLMsCharacter\-GLM\-6BDM\. vs RAG95\.6782\.35vs GraphRAG93\.1678\.59vs GCA85\.8369\.95CoSER\-8BDM\. vs RAG91\.4572\.67vs GraphRAG88\.6766\.90vs GCA89\.3556\.33
### 4\.2\.Experimental Settings

BaselinesWe compare DREAM against four representative baselines: \(1\)RAG\(Lewiset al\.,[2020](https://arxiv.org/html/2608.05170#bib.bib28)\): Retrieves text chunks most semantically similar to the scenario and incorporates them into role\-playing prompts for LLM response generation\. \(2\)GraphRAG\(Edgeet al\.,[2024](https://arxiv.org/html/2608.05170#bib.bib16)\): Retrieves nodes and community\-level summaries most relevant to the scenario from an indexed entity–relationship knowledge graph, and simulates role\-playing by combining task descriptions with retrieved information as prompts\. \(3\)Given\-Circumstance Acting \(GCA\)\(Wanget al\.,[2025](https://arxiv.org/html/2608.05170#bib.bib10)\): Utilizes LLMs to construct a book\-based global static profile containing rich character information, including persona, plots, experiences, and motivations, and simulates conversations based on the profile and scenario\. \(4\)CHARMAP\(Xuet al\.,[2025b](https://arxiv.org/html/2608.05170#bib.bib8)\): Retrieves scenario\-specific memories via vector matching and combines them with character profiles for persona\-driven decision\-making\.

DatasetWe evaluate DREAM on three datasets: 1\)CoSER\(Wanget al\.,[2025](https://arxiv.org/html/2608.05170#bib.bib10)\): We randomly select 100 books from the test set to evaluate general role\-playing ability via dialogue simulation; 2\)LIFECHOICE\(Xuet al\.,[2025b](https://arxiv.org/html/2608.05170#bib.bib8)\): A dataset focusing on RPAs’ decision\-making ability in multiple\-choice questions, where we randomly select 100 samples for evaluation; 3\)TCM: We identify the top 100 books on Goodreads’ Best Books Ever list111[https://www\.goodreads\.com/list/show/1\.Best\_Books\_Ever](https://www.goodreads.com/list/show/1.Best_Books_Ever), and select 20 books covering various genres as the evaluation dataset for TCM\. By constructing specific scenarios that can detect future knowledge leakage and causal narrative ability, to test the coherent temporal memory and long\-term causal narrative consistency of RPAs\.

ModelsOur experiments cover numerous LLMs: 1\) Closed\-source models, including GPT\-4o\(Hurstet al\.,[2024](https://arxiv.org/html/2608.05170#bib.bib29)\)and Gemini\-2\.5\-Pro\(Comaniciet al\.,[2025](https://arxiv.org/html/2608.05170#bib.bib30)\); 2\) Open\-source models, including LLaMA\-3\.1\-8B\-Instruct\(Grattafioriet al\.,[2024](https://arxiv.org/html/2608.05170#bib.bib31)\)and Qwen\-2\.5\-72B\-Instruct\(Yanget al\.,[2024](https://arxiv.org/html/2608.05170#bib.bib33)\); and 3\) Role\-playing models, including CoSER\-8B\(Wanget al\.,[2025](https://arxiv.org/html/2608.05170#bib.bib10)\)and CharacterGLM\-6B\(Zhouet al\.,[2024](https://arxiv.org/html/2608.05170#bib.bib1)\)\.

Implementation DetailsFor each benchmark, following the method outlined in §[3\.2\.3](https://arxiv.org/html/2608.05170#S3.SS2.SSS3), we construct EMG with an average of 32 events and 4,087 nodes per graph, and generate 30 challenging evaluation scenarios for each book\. We use LLM\-as\-a\-judge to conduct pairwise comparisons of dialogues generated by different methods across TCM metrics\(Chenet al\.,[2024b](https://arxiv.org/html/2608.05170#bib.bib22)\)\. We construct an independent EMG for each test role in benchmarks\. Implementation Details see Appendix[B](https://arxiv.org/html/2608.05170#A2)\.

All LLM\-based evaluations are conducted using GPT\-4o as the judge\. We verified the credibility of LLM\-as\-a\-judge through human evaluation by randomly selecting 60 scenarios from the 600 TCM test cases\. Results are presented in Appendix[D](https://arxiv.org/html/2608.05170#A4)\.

### 4\.3\.Main Results

##### Results on CoSER Benchmark

We evaluate DREAM alongside baseline models on the CoSER benchmark to assess general role\-playing capabilities\. Results, averaged over 10 runs, are presented in Table[1](https://arxiv.org/html/2608.05170#S3.T1)\. Benefiting from event\-based character profiles and structured knowledge captured in memory, DREAM consistently outperforms baseline models in storyline consistency, achieving a 14\.50% improvement and significantly reducing semantic drift in long\-term interactions\. With the Event Memory Graph \(EMG\) featuring causal narratives, it not only ensures that each response is consistent with the character in terms of style but also logically based on the causal cognitive chain\. Therefore, DREAM excels in storyline quality, achieving a 14\.88% improvement and effectively guaranteeing the story generation ability of LLMs in role\-playing tasks\. In addition, by leveraging dynamic character profiles, DREAM is almost on par with professional role\-playing models in terms of anthropomorphism and character fidelity\.

##### Results on LIFECHOICE Benchmark

We evaluate DREAM and other methods on LIFECHOICE for RPAs based on multi\-choice questions\. As shown in Table[2](https://arxiv.org/html/2608.05170#S3.T2), on the professional role\-playing model CoSER, the accuracy score of DREAM has increased by approximately 1\.21 points, and on the general model GPT\-4o, DREAM has also achieved an improvement of approximately 1\.82 points\. Benefiting from the causal chain and ability to prevent future knowledge leakage brought by DREAM, the model can better alleviate the problems of hallucinations and semantic drift in role\-playing\. It is worth noting that even when adding the motivations provided by the Benchmark, DREAM still maintains a stable performance advantage over closed\-source and open\-source models, which reflects the robustness and generalization ability of this method in role\-playing and memory\-related tasks\.

##### Results on TCM Benchmark

Table[3](https://arxiv.org/html/2608.05170#S4.T3)presents the winning rates of DREAM compared to the baselines in TCM evaluation\. Compared with RAG and GraphRAG, DREAM maintains a significant advantage in terms of Future Knowledge Leakage capability because it adds timestamps to character memories, with a win rate of over 90% on Closed\-Source and Role\-Playing LLMs\. Even when limited by the relatively small parameter size of Open\-Source LLMs, DREAM still achieves a win rate of over 80%\. Moreover, even a relatively comprehensive role\-playing method like GCA, in the absence of EMG, performs far worse than DREAM\. In terms of Causal Consistency capability, DREAM achieves an overall win rate of approximately 72%\. Therefore, it can be seen that incorporating a dynamically retrieved causal cognitive chain into the role\-playing process can provide a basis for character behaviors in long\-term and evolving role\-playing, avoid superficial style imitation, and significantly reduce the problem of model hallucinations\.

Table 4\.Ablation study results \(average scores\) on CoSER Test\. w/o T\.S\., w/o M\.R\. and w/o D\.P\. refer to removing timestamp, memory retrieval and dynamic profile, respectively\.ModelCompletew/o T\.S\.w/o M\.R\.w/o D\.P\.GPT\-4o62\.8658\.1857\.6759\.37Gemini\-2\.5\-Pro59\.4457\.0357\.3357\.16LLaMA\-3\.1\-8B57\.2254\.9554\.1655\.34Qwen\-2\.5\-72B59\.7658\.1757\.4558\.05CharacterGLM\-6B53\.7953\.1553\.2352\.94CoSER\-8B58\.0656\.9756\.3557\.17

### 4\.4\.Ablation Study

We conducted ablation studies on the main features of DREAM as shown in Table[4](https://arxiv.org/html/2608.05170#S4.T4)\. That is, we design three variants of DREAM by removing timestamp\(w/o T\.S\.\), dynamic profile\(w/o D\.P\.\) and the memory retrieval capability\(w/o M\.R\.\)\. This comparison was conducted on the CoSER benchmark, aiming to measure the role\-playing ability after the lack of key capabilities\.

Memory Retrieval \(M\.R\.\) has the largest impact\. Removing it leads to the most substantial performance drop, approximately 8\.9%, highlighting the central role of the constructed EMG in supporting role\-playing capabilities\. Notably, for high\-capacity models like GPT\-4o, the score plummets from62\.8662\.86to57\.6757\.67\. This underscores that the ability to retrieve causally relevant anchor events is indispensable for maintaining narrative depth in complex role\-playing\.

Timestamps \(T\.S\.\): The integration of timestamps is essential for preventing future knowledge leakage \(spoilers\)\. Without T\.S\., models consistently show a performance decline\. This suggests that while LLMs possess strong reasoning capabilities, they require explicit temporal constraints to correctly order retrieved memories\.

Dynamic Profile\(D\.P\.\) ensures that character cognitions and motivations evolve alongside the narrative\. Its removal results in a reduction about 5\.4% in character consistency\.

In summary, the synergy between the EMG\-based retrieval, temporal grounding, and dynamic profiling allows DREAM to achieve superior role\-playing fidelity that cannot be matched by any single component in isolation\.

## 5\.Conclusions

This paper presents DREAM, a novel framework for role\-playing that grounds dynamic character profiles in temporally and causally structured event\-centric memory\. DREAM transforms static literary texts into Event\-Aware Memory Graph \(EMG\), from which dual\-granularity retrievable character profiles are generated and updated via memory retrieval for event\-aware role\-playing\. In addition, we introduce the Temporal Causal Memory \(TCM\) benchmark to evaluate RPAs’ temporal consistency, memory accuracy, and causal narrative ability, complementing existing role\-playing evaluations\. Experiments show that DREAM, without task\-specific fine\-tuning, achieves state\-of\-the\-art results among all evaluated baselines across our proposed TCM benchmark and two existing RPA benchmarks, outperforming existing role\-playing models\. Our approach demonstrates the effectiveness of structured memory in enhancing the interpretability and fidelity of role\-playing agents\. Future work will explore memory mechanisms that incrementally integrate interaction\-generated dialogues into the EMG, enabling agents to evolve continuously\.

## Limitations

Despite DREAM’s strong performance on CoSER and TCM benchmarks, it has limitations: Role\-playing fidelity depends on extraction precision—EMG preserves core narrative logic but may abstract fine\-grained stylistic details, leaving room for more expressive future modeling\. Though memory construction is a one\-time offline process, graph\-based multi\-hop retrieval and dynamic profile generation incur higher inference latency and token costs than standard vector\-based RAG\. Consistent with common benchmarking standards, we use LLM\-based judges\(Zhenget al\.,[2023](https://arxiv.org/html/2608.05170#bib.bib34); Liet al\.,[2025](https://arxiv.org/html/2608.05170#bib.bib35)\)\. While aligned with human judgment, these metrics are limited by automated assessment inherent traits, a standard consideration in current narrative evaluation\.

## Ethics Statement

In this paper, we introduce the DREAM framework\. The design and utilization of this framework are strictly guided by ethical principles to ensure its applications yield beneficial outcomes for society\. We have curated and extracted sample data from a diverse and representative set of literary novels and plays from both domestic and international sources; these texts are utilized exclusively for the construction and evaluation of natural language processing models, with the objective of advancing scientific research in this field\. While our framework is designed to enhance role\-playing systems, dialogue generation still retains a degree of unpredictability\. Consequently, in high\-stakes or sensitive scenarios, it is essential to perform rigorous safety scrutiny on generated responses to ensure their appropriateness\. We encourage the responsible use of DREAM for educational, entertainment, and creative purposes, while discouraging any harmful or malicious activities\.

## References

- J\. Ahn, T\. Lee, J\. Lim, J\. Kim, S\. Yun, H\. Lee, and G\. Kim \(2024\)Timechara: evaluating point\-in\-time character hallucination of role\-playing large language models\.arXiv preprint arXiv:2405\.18027\.Cited by:[§1](https://arxiv.org/html/2608.05170#S1.p2.1)\.
- J\. Chen, X\. Wang, R\. Xu, S\. Yuan, Y\. Zhang, W\. Shi, J\. Xie, S\. Li, R\. Yang, T\. Zhu,et al\.\(2024a\)From persona to personalization: a survey on role\-playing language agents\.arXiv preprint arXiv:2404\.18231\.Cited by:[§1](https://arxiv.org/html/2608.05170#S1.p1.1)\.
- J\. Chen, X\. Zhu, C\. Yang, C\. Shi, Y\. Xi, Y\. Zhang, J\. Wang, J\. Pu, T\. Feng, Y\. Yang,et al\.\(2024b\)Hollmwood: unleashing the creativity of large language models in screenwriting via role playing\.InFindings of the Association for Computational Linguistics: EMNLP 2024,pp\. 8075–8121\.Cited by:[§3\.3\.2](https://arxiv.org/html/2608.05170#S3.SS3.SSS2.p2.1),[§4\.2](https://arxiv.org/html/2608.05170#S4.SS2.p4.1)\.
- N\. Chen, Y\. Wang, H\. Jiang, D\. Cai, Y\. Li, Z\. Chen, L\. Wang, and J\. Li \(2023\)Large language models meet harry potter: a dataset for aligning dialogue agents with characters\.InFindings of the association for computational linguistics: EMNLP 2023,pp\. 8506–8520\.Cited by:[§2\.2](https://arxiv.org/html/2608.05170#S2.SS2.p1.1)\.
- G\. Comanici, E\. Bieber, M\. Schaekermann, I\. Pasupat, N\. Sachdeva, I\. Dhillon, M\. Blistein, O\. Ram, D\. Zhang, E\. Rosen,et al\.\(2025\)Gemini 2\.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities\.arXiv preprint arXiv:2507\.06261\.Cited by:[§4\.2](https://arxiv.org/html/2608.05170#S4.SS2.p3.1)\.
- Y\. Dai, H\. Hu, L\. Wang, S\. Jin, X\. Chen, and Z\. Lu \(2024\)Mmrole: a comprehensive framework for developing and evaluating multimodal role\-playing agents\.arXiv preprint arXiv:2408\.04203\.Cited by:[§2\.1](https://arxiv.org/html/2608.05170#S2.SS1.p1.1)\.
- D\. Edge, H\. Trinh, N\. Cheng, J\. Bradley, A\. Chao, A\. Mody, S\. Truitt, D\. Metropolitansky, R\. O\. Ness, and J\. Larson \(2024\)From local to global: a graph rag approach to query\-focused summarization\.arXiv preprint arXiv:2404\.16130\.Cited by:[§2\.2](https://arxiv.org/html/2608.05170#S2.SS2.p1.1),[§4\.2](https://arxiv.org/html/2608.05170#S4.SS2.p1.1)\.
- A\. Ellis \(1957\)Rational psychotherapy and individual psychology\.Journal of individual psychology13\(1\),pp\. 38\.Cited by:[§1](https://arxiv.org/html/2608.05170#S1.p4.1)\.
- J\. Fang, H\. Jiang, K\. Wang, Y\. Ma, S\. Jie, X\. Wang, X\. He, and T\. Chua \(2024\)Alphaedit: null\-space constrained knowledge editing for language models\.arXiv preprint arXiv:2410\.02355\.Cited by:[§2\.2](https://arxiv.org/html/2608.05170#S2.SS2.p1.1)\.
- A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan,et al\.\(2024\)The llama 3 herd of models\.arXiv preprint arXiv:2407\.21783\.Cited by:[§4\.2](https://arxiv.org/html/2608.05170#S4.SS2.p3.1)\.
- K\. He, Y\. Huang, W\. Wang, D\. Ran, D\. Sheng, J\. Huang, Q\. Lin, J\. Xu, W\. Liu, and M\. Feng \(2025\)Crab: a novel configurable role\-playing llm with assessing benchmark\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 15030–15052\.Cited by:[§1](https://arxiv.org/html/2608.05170#S1.p2.1),[§2\.1](https://arxiv.org/html/2608.05170#S2.SS1.p1.1)\.
- Y\. Hu, S\. Liu, Y\. Yue, G\. Zhang, B\. Liu, F\. Zhu, J\. Lin, H\. Guo, S\. Dou, Z\. Xi,et al\.\(2025\)Memory in the age of ai agents\.arXiv preprint arXiv:2512\.13564\.Cited by:[§2\.2](https://arxiv.org/html/2608.05170#S2.SS2.p1.1)\.
- A\. Hurst, A\. Lerer, A\. P\. Goucher, A\. Perelman, A\. Ramesh, A\. Clark, A\. Ostrow, A\. Welihinda, A\. Hayes, A\. Radford,et al\.\(2024\)Gpt\-4o system card\.arXiv preprint arXiv:2410\.21276\.Cited by:[§4\.2](https://arxiv.org/html/2608.05170#S4.SS2.p3.1)\.
- P\. Lewis, E\. Perez, A\. Piktus, F\. Petroni, V\. Karpukhin, N\. Goyal, H\. Küttler, M\. Lewis, W\. Yih, T\. Rocktäschel,et al\.\(2020\)Retrieval\-augmented generation for knowledge\-intensive nlp tasks\.Advances in neural information processing systems33,pp\. 9459–9474\.Cited by:[§4\.2](https://arxiv.org/html/2608.05170#S4.SS2.p1.1)\.
- C\. Li, Z\. Leng, C\. Yan, J\. Shen, H\. Wang, W\. Mi, Y\. Fei, X\. Feng, S\. Yan, H\. Wang,et al\.\(2023\)Chatharuhi: reviving anime character in reality via large language model\.arXiv preprint arXiv:2308\.09597\.Cited by:[§1](https://arxiv.org/html/2608.05170#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.05170#S2.SS2.p1.1),[§3\.2\.2](https://arxiv.org/html/2608.05170#S3.SS2.SSS2.p1.1)\.
- D\. Li, B\. Jiang, L\. Huang, A\. Beigi, C\. Zhao, Z\. Tan, A\. Bhattacharjee, Y\. Jiang, C\. Chen, T\. Wu,et al\.\(2025\)From generation to judgment: opportunities and challenges of llm\-as\-a\-judge\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 2757–2791\.Cited by:[§4](https://arxiv.org/html/2608.05170#S4.p1.1),[Limitations](https://arxiv.org/html/2608.05170#Sx1.p1.1)\.
- C\. Liao, K\. Wang, Y\. Wu, F\. Huang, and Y\. Li \(2025\)MOA: multi\-objective alignment for role\-playing agents\.arXiv preprint arXiv:2512\.09756\.Cited by:[§2\.1](https://arxiv.org/html/2608.05170#S2.SS1.p1.1)\.
- C\. Liu, Y\. Lu, F\. Ye, J\. Li, X\. Chen, F\. Ren, Z\. Tu, and X\. Li \(2025\)CogDual: enhancing dual cognition of llms via reinforcement learning with implicit rule\-based rewards\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 27295–27324\.Cited by:[§2\.1](https://arxiv.org/html/2608.05170#S2.SS1.p1.1)\.
- K\. Lu, B\. Yu, C\. Zhou, and J\. Zhou \(2024\)Large language models are superpositions of all characters: attaining arbitrary role\-play via self\-alignment\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 7828–7840\.Cited by:[§1](https://arxiv.org/html/2608.05170#S1.p2.1),[§2\.1](https://arxiv.org/html/2608.05170#S2.SS1.p1.1)\.
- D\. Occhipinti, S\. S\. Tekiroğlu, and M\. Guerini \(2024\)Prodigy: a profile\-based dialogue generation dataset\.InFindings of the Association for Computational Linguistics: NAACL 2024,pp\. 3500–3514\.Cited by:[§2\.1](https://arxiv.org/html/2608.05170#S2.SS1.p1.1)\.
- J\. S\. Park, J\. O’Brien, C\. J\. Cai, M\. R\. Morris, P\. Liang, and M\. S\. Bernstein \(2023\)Generative agents: interactive simulacra of human behavior\.InProceedings of the 36th annual acm symposium on user interface software and technology,pp\. 1–22\.Cited by:[§1](https://arxiv.org/html/2608.05170#S1.p1.1)\.
- Y\. Tang, K\. Chen, M\. Yang, Z\. Niu, J\. Li, T\. Zhao, and M\. Zhang \(2025\)Thinking in character: advancing role\-playing agents with role\-aware reasoning\.arXiv preprint arXiv:2506\.01748\.Cited by:[§2\.1](https://arxiv.org/html/2608.05170#S2.SS1.p1.1)\.
- Y\. Tseng, Y\. Huang, T\. Hsiao, W\. Chen, C\. Huang, Y\. Meng, and Y\. Chen \(2024\)Two tales of persona in llms: a survey of role\-playing and personalization\.arXiv preprint arXiv:2406\.01171\.Cited by:[§1](https://arxiv.org/html/2608.05170#S1.p1.1)\.
- N\. Wang, Z\. Peng, H\. Que, J\. Liu, W\. Zhou, Y\. Wu, H\. Guo, R\. Gan, Z\. Ni, J\. Yang,et al\.\(2024\)Rolellm: benchmarking, eliciting, and enhancing role\-playing abilities of large language models\.InFindings of the Association for Computational Linguistics: ACL 2024,pp\. 14743–14777\.Cited by:[§1](https://arxiv.org/html/2608.05170#S1.p1.1),[§2\.2](https://arxiv.org/html/2608.05170#S2.SS2.p1.1)\.
- X\. Wang, H\. Wang, Y\. Zhang, X\. Yuan, R\. Xu, J\. Huang, S\. Yuan, H\. Guo, J\. Chen, S\. Zhou,et al\.\(2025\)Coser: coordinating llm\-based persona simulation of established roles\.arXiv preprint arXiv:2502\.09082\.Cited by:[Appendix C](https://arxiv.org/html/2608.05170#A3.p1.1),[§1](https://arxiv.org/html/2608.05170#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.05170#S2.SS2.p1.1),[§4\.1](https://arxiv.org/html/2608.05170#S4.SS1.p1.1),[§4\.2](https://arxiv.org/html/2608.05170#S4.SS2.p1.1),[§4\.2](https://arxiv.org/html/2608.05170#S4.SS2.p2.1),[§4\.2](https://arxiv.org/html/2608.05170#S4.SS2.p3.1)\.
- B\. Xu, S\. Zhao, R\. Wu, Z\. Huang, J\. Wang, Z\. Hu, K\. Wang, H\. Liu, T\. Lv, L\. Li,et al\.\(2025a\)Empowering economic simulation for massively multiplayer online games through generative agent\-based modeling\.InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V\. 2,pp\. 3366–3377\.Cited by:[§1](https://arxiv.org/html/2608.05170#S1.p1.1)\.
- R\. Xu, X\. Wang, J\. Chen, S\. Yuan, X\. Yuan, J\. Liang, Z\. Chen, Y\. Xiao,et al\.\(2025b\)Character is destiny: can persona\-assigned language models make personal choices?\.InFindings of the Association for Computational Linguistics: EMNLP 2025,pp\. 15038–15059\.Cited by:[§1](https://arxiv.org/html/2608.05170#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.05170#S2.SS2.p1.1),[§3\.2\.2](https://arxiv.org/html/2608.05170#S3.SS2.SSS2.p1.1),[§4\.2](https://arxiv.org/html/2608.05170#S4.SS2.p1.1),[§4\.2](https://arxiv.org/html/2608.05170#S4.SS2.p2.1)\.
- S\. Yan, X\. Yang, Z\. Huang, E\. Nie, Z\. Ding, Z\. Li, X\. Ma, K\. Kersting, J\. Z\. Pan, H\. Schütze,et al\.\(2025\)Memory\-r1: enhancing large language model agents to manage and utilize memories via reinforcement learning\.arXiv preprint arXiv:2508\.19828\.Cited by:[§2\.2](https://arxiv.org/html/2608.05170#S2.SS2.p1.1)\.
- A\. Yang, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Li, D\. Liu, F\. Huang, G\. Dong, H\. Wei, H\. Lin, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Lin, K\. Dang, K\. Lu, K\. Bao, K\. Yang, L\. Yu, M\. Li, M\. Xue, P\. Zhang, Q\. Zhu, R\. Men, R\. Lin, T\. Li, T\. Xia, X\. Ren, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Cui, Z\. Zhang, Z\. Qiu, S\. Quan, and Z\. Wang \(2024\)Qwen2\.5 technical report\.arXiv preprint arXiv:2412\.15115\.Cited by:[§4\.2](https://arxiv.org/html/2608.05170#S4.SS2.p3.1)\.
- X\. Ye, R\. Wang, Y\. Wu, V\. Ma, F\. Fang, F\. Huang, and Y\. Li \(2025\)Cpo: addressing reward ambiguity in role\-playing dialogue via comparative policy optimization\.Preprint\.Cited by:[§2\.1](https://arxiv.org/html/2608.05170#S2.SS1.p1.1)\.
- H\. Zhang, R\. Luo, X\. Liu, Y\. Wu, T\. Lin, P\. Zeng, Q\. Qu, F\. Fang, M\. Yang, L\. Gao,et al\.\(2025a\)OmniCharacter: towards immersive role\-playing agents with seamless speech\-language personality interaction\.arXiv preprint arXiv:2505\.20277\.Cited by:[§2\.1](https://arxiv.org/html/2608.05170#S2.SS1.p1.1)\.
- Z\. Zhang, Q\. Dai, X\. Bo, C\. Ma, R\. Li, X\. Chen, J\. Zhu, Z\. Dong, and J\. Wen \(2025b\)A survey on the memory mechanism of large language model\-based agents\.ACM Transactions on Information Systems43\(6\),pp\. 1–47\.Cited by:[§2\.2](https://arxiv.org/html/2608.05170#S2.SS2.p1.1)\.
- L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. Xing,et al\.\(2023\)Judging llm\-as\-a\-judge with mt\-bench and chatbot arena\.Advances in neural information processing systems36,pp\. 46595–46623\.Cited by:[§4](https://arxiv.org/html/2608.05170#S4.p1.1),[Limitations](https://arxiv.org/html/2608.05170#Sx1.p1.1)\.
- J\. Zhou, Z\. Chen, D\. Wan, B\. Wen, Y\. Song, J\. Yu, Y\. Huang, P\. Ke, G\. Bi, L\. Peng,et al\.\(2024\)CharacterGLM: customizing social characters with large language models\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track,pp\. 1457–1476\.Cited by:[§1](https://arxiv.org/html/2608.05170#S1.p1.1),[§1](https://arxiv.org/html/2608.05170#S1.p2.1),[§2\.1](https://arxiv.org/html/2608.05170#S2.SS1.p1.1),[§4\.2](https://arxiv.org/html/2608.05170#S4.SS2.p3.1)\.
- X\. Zhou, H\. Zhu, L\. Mathur, R\. Zhang, H\. Yu, Z\. Qi, L\. Morency, Y\. Bisk, D\. Fried, G\. Neubig,et al\.\(2023\)Sotopia: interactive evaluation for social intelligence in language agents\.arXiv preprint arXiv:2310\.11667\.Cited by:[§1](https://arxiv.org/html/2608.05170#S1.p1.1)\.

## Appendix APrompts and KG Definitions for Macro and Micro\-Level Extraction

In this section, we provide the detailed ontology definitions and extraction prompts used in our Dual\-Granularity Knowledge Extraction module§[3\.2\.2](https://arxiv.org/html/2608.05170#S3.SS2.SSS2)\.

### A\.1\.Knowledge Graph Schema Definition

To capture both the global context and fine\-grained narrative dynamics, we designed two distinct schemas for the Event\-Aware Memory Graph \(EMG\)\.

##### Macro\-Level Event Context\.

As mentioned in Section §[3\.2\.2](https://arxiv.org/html/2608.05170#S3.SS2.SSS2), we define 8 types of entities and 7 types of relations to capture the static context of an eventEiE\_\{i\}\. Table[6](https://arxiv.org/html/2608.05170#A5.T6)details the definitions of these schema elements, covering world settings, background information, and character attributes\.

##### Micro\-Level Dynamic Narrative\.

To model the causal evolution withinUpU\_\{p\}, we define 8 types of entities and 12 types of relations\. As shown in Table[7](https://arxiv.org/html/2608.05170#A5.T7), these definitions focus on the Emotion\-Cognition\-Behavior chains, explicitly representing the logical flow of character development\.

## Appendix BImplementation Details

In CoSER, we replaced the original profile with the DREAM profile generated based on the scenario\. In LIFECHOICE and TCM, we use the complete DREAM method, which includes a dynamic profile and memory retrieval\.

We maintain the number of dialogue turns and scenarios consistent across all methods\. Each experiment involves 2–4 simulated scenarios, and each scenario\-method pair is run 10 times\. The highest and lowest scores are discarded, and the mean of the remaining runs is taken as the final score\.

### B\.1\.Extraction Prompts

We employ Large Language Models \(LLMs\) to extract structured information from raw literary texts\. The prompts are designed to ensure consistency in entity naming and coreference resolution\. Table[8](https://arxiv.org/html/2608.05170#A5.T8)and Table[9](https://arxiv.org/html/2608.05170#A5.T9)present the specific instructions used for Macro\-Level and Micro\-Level extraction, respectively\.

## Appendix CDetailed Evaluation Metrics

Following the methodology of CoSER\(Wanget al\.,[2025](https://arxiv.org/html/2608.05170#bib.bib10)\), we employ GPT\-4o as a critic to evaluate simulated role\-playing conversations\. The detailed definitions for the four key dimensions are as follows:

\(1\)Storyline Consistency:Assesses alignment between simulated conversations and original dialogue, focusing on whether RPAs’ responses \(emotions, attitudes, behaviors\) remain faithful to the narrative context\.

\(2\)Anthropomorphism:Assesses if RPAs act human\-like via rubrics covering self\-identity, emotional depth, persona coherence, and social interaction\.

\(3\)Character Fidelity:Assesses how well RPAs reflect character, including linguistic style, knowledge and background, personality, behavior, and relationships\.

\(4\)Storyline Quality:Assesses the naturalness of simulated conversations, focusing on narrative flow and logical consistency\.

## Appendix DConsistency with Human Evaluation

To validate the reliability of our model\-based evaluation approach used in the TCM benchmark \(§[4\.2](https://arxiv.org/html/2608.05170#S4.SS2)\), we conducted a comprehensive agreement analysis between the LLM judge \(GPT\-4o\) and human assessments\.

##### Setup\.

We recruited 5 human annotators who are avid readers and “fans” of the corresponding literary works to ensure they possess the necessary domain knowledge to judge character behavior and plot consistency\. As mentioned in Section \(§[4\.2](https://arxiv.org/html/2608.05170#S4.SS2)\), we randomly sampled 60 evaluation scenarios from the total 600 test scenarios in the TCM benchmark\. For each scenario, annotators were presented with pairwise outputs \(DREAM vs\. Baseline\) and asked to determine the winner \(or tie\) across two dimensions:Future Knowledge Leakage \(FKL\.\)andCausal Consistency \(CC\.\)\.

##### Agreement Analysis\.

To quantify the inter\-rater reliability between the model’s judgments and human annotations, we utilized Cohen’s Kappa coefficient \(κ\\kappa\), which accounts for the possibility of chance agreement\. The results are presented in Table[5](https://arxiv.org/html/2608.05170#A4.T5)\. The analysis reveals Cohen’s Kappa scores ranging from 0\.68 to 0\.79 across the three metrics\. Specifically,FKL\.achieved the highest agreement \(κ=0\.786\\kappa=0\.786\), likely because temporal contradictions \(e\.g\., mentioning future events\) are objectively verifiable\.CC\.also demonstrated substantial agreement \(κ\>0\.65\\kappa\>0\.65\)\. These results indicate a high level of consistency between human and machine evaluations, confirming that our LLM\-as\-a\-judge paradigm is sufficiently reliable for reflecting the actual performance differences in the TCM benchmark\.

MetricFKL\.CC\.Cohen’s Kappa \(κ\\kappa\)0\.7860\.688Table 5\.The Cohen’s Kappa \(κ\\kappa\) agreement between Human Evaluation and LLM\-based Evaluation on TCM metrics\.

## Appendix EPrompts Demostration

We provide the details of the prompt templates of DREAM in this section\.

The prompt for incremental character profile updating and refinement is displayed in Table[11](https://arxiv.org/html/2608.05170#A5.T11)\. The prompt for character\-centric event recognition and classification is displayed in Table[10](https://arxiv.org/html/2608.05170#A5.T10)\.

Table 6\.Entity and Relation Definitions for Macro\-Level Event Context extraction\.CategoryNameDescriptionEntityWorld SettingThe worldview settings of the current event\.EnvironmentThe description and characteristics of the environmental scene\.DialogueThe dialogue or inner monologue of the core character\.RelationThe relationship between the selected character and other characters\.IdentityThe identity of the selected character\.CharacterOther characters related to the core character in the event\.AppearanceThe description of the selected character’s appearance\.ActionSpecific actions or habits of the selected character\.Relationhas\_worldviewLinks an event to its corresponding World Setting\.has\_envLinks an event to its environmental description\.has\_relation\_withRelationships between characters\.has\_identityDescribes the role or identity of selected character in the specific event\.has\_dialogueLinks a dialogue to selected character\.has\_actionLinks a specific action to selected character\.has\_appearanceLinks the appearance description to the selected character\.Table 7\.Entity and Relation Definitions for Micro\-Level Dynamic Narrative extraction\.CategoryNameDescriptionEntityUnitPlotSub\-plot units that constitute the selected characteristics of an event\.CharacterOther characters related to the selected character in the event\.EmotionSpecific emotions of the selected character in the Sub\-plot\.CognitionCognitive views formed/updated by the selected character in the Sub\-plot\.BehaviorKey behavioral actions in the Sub\-plot that promote the development of the event\.SceneSpecific scene/location where the event occurs\.ItemInteractive and owned objects of the selected character’s behaviors\.SkillSpecial skills or abilities of the selected character\.Relationnext\_plotDefines the temporal sequence between unit\-plots\.core\_roleThe core character of the sub\-plot\.other\_roleOther participants in the Sub\-plot\.has\_behaviorCore behaviors of Role in the Sub\-plot\.interact\_objectInteractive objects \(item or skill\) in the behavior\.occurs\_atAssociation with the scene where the plot occurs\.emotion\_fromTriggering cause of character’s emotion\.cognition\_fromSource of formation of character’s cognition\.cognition\_update\_toChanges in character’s cognition\.motivated\_byDriving factors of character’s behavior\.prefer\_toSpecific objects of character’s likes and preferences\.prefer\_causeReasons for preferences\.Table 8\.Prompt for Macro\-Level Event Context Extraction\.\# Your Role: An assistant adept at mining information based on the event content and requirements of the book “\{BookName\}” provided by users

\# Task Requirements 1\. Users will provide the content text corresponding to the events for extraction\. 2\. Complete the following tasks and conduct extraction as richly and meaningfully as possible\. 3\. The use of personal pronouns such as “you”, “I”, “he”, “she” is prohibited; descriptions must use the clear and unified character names mentioned in the event summary\.

\# Task 1: Extract Content by Predefined Categories 1\. WorldSetting: The worldview settings of the current event 2\. Environment: The description and characteristics of the environmental scene of the current event 3\. Dialogue: The ‘Dialogue or Inner monologue’ of the core character \{Role\}, as well as the corresponding linguistic style \(tone, commonly used vocabulary, expression habits\) 4\. Relation: The relationship between the core character \{Role\} and other characters \(e\.g\., friend, foe, leader, follower, etc\.\) 5\. Identity: The identity of the core character \(titles, nicknames, aliases, and other designations that can refer to \{Role\}\) 6\. Character: Other characters related to the core character in the event 7\. Appearance: The description of the core character \{Role\}’s appearance in the event 8\. Action: Specific actions or habits of the core character \{Role\} that can reflect the character’s personality, emotions, persona, representativeness, and iconicity \(e\.g\., pushing up glasses when thinking, preferring to speak with a soft chuckle when acting as a deity, speaking with a sigh when helpless, being patient to explain when elaborating, habitually roaring when angry……\)

\# Important Notes: \- Add as rich and meaningful attributes as possible during the content extraction process to facilitate the subsequent construction of knowledge graph triples for the event\. \- If there is no corresponding content for a category during extraction, ignore that category\. \- Extract entities in the order they appear in the text\.

\# Initial Settings Your role positioning: An assistant adept at extracting information based on the event content and requirements of the book “\{BookName\}” provided by users\. Strictly comply with the above specific task requirements and output results in \{Language\}\.

Table 9\.Prompt for Micro\-Level Dynamic Narrative Extraction\.\# Your role: Event Knowledge Graph Constructor based on the content of the book \{BookName\} Based on the complete event text, with the specified core character \{Role\} as the center, systematically sort out the development context of Sub\-plot units, emotional dynamic changes, cognitive iteration processes, and behavioral logical associations in the event content\. Finally, generate an Event Knowledge Graph adapted to role\-playing scenarios, providing a basis for the Role Memory Module and Behavioral Decision\-making\.

\# Core Task Objectives 1\. With the core character \{Role\} as the center, disassemble the large event content into logically connected Sub\-plot units to form a traceable event chain\. 2\. Summarize the emotional fluctuations \[including triggering reasons and lasting impacts\], cognitive changes \[including cognitive sources and subsequent effects\] of the core character \{Role\} in the Sub\-plots, and associate the physical elements of the events \[behaviors, scenes, objects\]\. 3\. The constructed graph must fit the needs of role\-playing, which can not only restore the original appearance of the event but also support the role to replicate the corresponding emotional state, follow cognitive logic, and echo past events in subsequent interactions\. 4\. Prohibit the use of pronouns such as you, I, he, she, etc\. Must use the clear role names mentioned in the event summary for description, for example: \{Role\}\. 5\. The final output is in the JSON format of the example

\# Ontology Definition \*\*1\. Nodes \[Entities\]:\*\* 1\. UnitPlot: Sub\-plot units that constitute the core characteristics of an event 2\. Character: All roles in the event, the core character is fixed as \{Role\}, “other\_role” are labeled with specific role names 3\. Emotion: Specific emotions of the core character in the Sub\-plot 4\. Cognition: Cognitive views formed/updated by the core character in the Sub\-plot \[need to label the type of cognition: Basic Cognition/Value Judgment/Behavioral Norm\] 5\. Behavior: Key behavioral actions of the core character in the Sub\-plot that promote the development of the event 6\. Scene: Specific scene/location where the event occurs 7\. Item: Interactive and owned objects of the core character \{Role\}’s behaviors 8\. Skill: Special skills and abilities of the core character \{Role\}

\*\*2\. Relationships \[Edges\]:\*\* 1\. next\_plot: Temporal sequence association between small plots \[“UnitPlot1”, “next\_event”, “UnitPlot2”\]\. 2\. core\_role: Core leader \{Role\} of the Sub\-plot \[“UnitPlot”, “core\_role”, “Role”\] 3\. other\_role: Other participants in the Sub\-plot \[“UnitPlot”, “other\_role”, “Role”\] 4\. has\_behavior: Core behaviors of \{Role\} in the Sub\-plot \[“UnitPlot”, “has\_behavior”, “Behavior”\] 5\. interact\_object: Interactive objects in the behavior \[“UnitPlot”, “interact\_object”, “Object”\] 6\. occurs\_at: Association with the scene where the plot occurs \[“UnitPlot”, “occurs\_at”, “Scene”\] 7\. emotion\_from: Triggering cause of \{Role\}’s emotion \[“Emotion”, “emotion\_from”, “Summary for forming the emotion”\] 8\. cognition\_from: Source of formation of \{Role\}’s cognition \[“Cognition”, “cognition\_from”, “Summary of the reasons for forming the cognition”\] 9\. cognition\_update\_to: Changes in \{Role\}’s cognition \[“old\_cognition”, “cognition\_update\_to”, “new\_cognition”\] 10\. motivated\_by: Driving factors of \{Role\}’s behavior \[“Behavior”, “motivated\_by”, “Emotion/Cognition”\] 11\. prefer\_to: Specific objects of \{Role\}’s likes and preferences \[“Sub\-plot X”, “prefer\_to”, “Object”\] 12\. prefer\_cause: Reasons for preferences \[“Object”, “prefer\_cause”, “Cognition”\]

\# Task Execution Requirements 1\. Input Adaptation: Based on the large event content provided by the user, conduct structured disassembly within the existing event framework\. 2\. Core Anchoring: The sorting out of all Sub\-plots, emotions, and cognition must revolve around \{Role\}, prioritizing the presentation of the event perception and psychological changes from their perspective\. 3\. Format Specification: Directly output the structured JSON list of “unit\_events” in the reference example, without wrapping the final output content with “‘json\. 4\. Strictly abide by the above specific task requirements, and output a structured json list in \{Language\}, which conforms to the json structure in the example\.

\# User Input Text Example: \{User\_input\_Example\} \# Example of structured output in final json format:\{out\_put\_example\}

Table 10\.Prompt for Event Recognition and Classification\.\# Your Role: An event understanding assistant proficient in the book \{BookName\}, who combines the user\-provided single chapter content and previous plot summaries

\#\# Task 1: Summarize Content Type 1\. Define the candidate output types as: \[“complete event”, “incomplete event”, “non\-character event”, “contextual content”\] 2\. Definition scope and explanation of types: \- A “complete event” refers to the content of the chapter that involves event of the role \{Role\}, and the content of the current chapter completely describes the event\. \- An “incomplete event” refers to an event involving the character \{Role\}, but the chapter content provided by the user is insufficient to fully describe the current event, or the current event has not ended and requires additional content from subsequent chapters\. \- A “non\-character event” refers to an event in the current chapter content that involves other characters, does not directly include the character \{Role\}, but the chapter content fully describes the event\. \- A “contextual content” refers to content in the current chapter that is not an event, such as background descriptions, previous plot summaries, introductions of character relationships, and other similar contextual content descriptions\. 3\. Based on the specific content of the chapter in \{BookName\} provided by the user and whether \{Role\} is the central figure, output a content type\. 4\. The summarized content type must strictly follow the defined array\. 5\. NOTE: Task 1 is only a preliminary task for Task 2 and Task 3, and should not be output separately in the end\.

\#\# Task 2: Determine Content Type 1\. If the content type output in \[Task 1\] is \[“complete event”\], do not output it first and continue to \[Task 3\] before outputting\. 2\. If the content type output in \[Task 1\] is \[“incomplete event”\], output strictly in accordance with the example array format and content: \[\{\{“event\_type”:“incomplete event”, “event\_time”:“time marker”, “event\_name”:“summarized name of the incomplete event”, “event\_description”:“summary of the content of the incomplete event chapter”\}\}\], and end all tasks\. 3\. If the content type output in \[Task 1\] is \[“non\-character event”\], output strictly in accordance with the example array format and content: \[\{\{“event\_type”:“non\-character event”, “event\_time”:“time marker”, “event\_name”:“summarized name of the non\-character event”, “event\_description”:“summary of the content of the non\-character event chapter”\}\}\], and end all tasks\. 4\. If the content type output in \[Task 1\] is \[“contextual content”\], output strictly in accordance with the example array format and content: \[\{\{“event\_type”:“contextual content”, “event\_time”:“time marker”, “event\_name”:“summarized name of the contextual content”, “event\_description”:“summary of the content of the contextual content chapter”\}\}\], and end all tasks\.

\#\# Task 3: Summarize and Extract Events and Requirements 1\. Take \{Role\} as the central figure of the event content\. 2\. Based on the specific content of the chapter in \{BookName\} provided by the user, summarize it into an event centered on \{Role\}\. 3\. According to the chapter content, summarize a time or time period that can mark the sequence of events in the book; if the time cannot be determined, use the chapter number as the time marker\. 4\. It is forbidden to use personal pronouns such as you, I, he, she, etc\. The event summary must use clear character names, character titles, or character identities, etc\. 5\. The output content and format are: \[\{\{“event\_type”:“complete event”, “event\_time”:“time marker”, “event\_name”:“summarized event name”, “event\_description”:“summarized and induced event content description”\}\}\]

\#\# Initialization You act as: An event understanding assistant proficient in the book \{BookName\}, who refers to previous plot summaries and combines the user\-provided single chapter content\. Strictly follow the order of tasks and output the results using \{Language\}\.

Table 11\.Prompt for Incremental Character Profile Updating and Refinement\.Role: You are an expert literary analyst and character profiler of the book \{BookName\}\. Your task is tocontinuously update and refinethe Character Profile for \{Role\} as the narrative progresses\.

Inputs: 1\.Existing Profile:The character’s profile derived from previous events \(if any\)\. 2\.New Event Data:Extracted knowledge graph data from the current event \{Timeline: T\_current\}, including Identity, Background, Description, Style, Traits, Motivations, Relationships, Causal Narrative and Cognition\.

Instructions: 1\. Incremental Update: DO NOT discard the “Existing Profile”\. Use the “New Event Data” to enrich, verify, or evolve the existing profile\. \-Reinforce:If new data confirms existing traits, strengthen the description\. \-Add:If new data reveals previously unknown aspects \(e\.g\., a new relationship or hidden skill\), add them to the relevant section\. \-Evolve:If the character undergoes a change \(e\.g\., from calm to angry, or a change in worldview\), explicitly describe this evolution in the “Cognitive State” or “Key Experiences” section\. \-Contextualize:Resolve conflicts based on the timeline\. If the character was “Loyal” in T1 but “Betrayed” in T2, the profile should reflect this shift\. 2\. Synthesize: Write in coherent, literary paragraphs\. Do not simply append lists\. Merge new information naturally into the existing structure\. 3\. Output Format: Strictly follow the section structure below\. Output the results using \{Language\}\.

Target Structure: \*\*Name:\*\*\{Role\} \*\*Background:\*\*\(Update based on new status or revealed backstory\) \*\*Appearance:\*\*\(Add details if appearance changes or new features are described\) \*\*Linguistic styles:\*\*\(Update if the character’s tone shifts in this event\) \*\*Personality Traits:\*\*\(Refine traits based on new actions\) \*\*Core Motivations:\*\*\(Update current drives and underlying motives\) \*\*Relationships:\*\*\(Update dynamic relations with others\) \*\*Cognition Chains:\*\*\(CRITICAL: Summarize the character’s state of mind in this event and how it compares to the past\) \*\*Key Experiences:\*\*\(Briefly add the essence of the current event to their history\) \*\*Causal Narrative Chains\*\*\(Generate a causal narrative paragraph description based on provided triples\)

Similar Articles

MemoryForge: Synthesize Lifelong Memory for Human-Like LLM Agents

arXiv cs.CL

This paper introduces MemoryForge, a framework for synthesizing lifelong autobiographical memory from brief target personas to enable frozen LLMs to exhibit more human-like behaviors in role-play and user-simulation, outperforming descriptive conditioning baselines.

AdMem: Advanced Memory for Task-solving Agents

arXiv cs.AI

This paper introduces AdMem, a unified memory framework for LLM-based agents that integrates semantic, episodic, and procedural memory with a bi-level short-term and long-term store, using a multi-agent architecture for automatic memory generation and adaptive retrieval. Experiments show improved robustness and success on long multi-turn tasks.

Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability

arXiv cs.CL

This paper presents the first systematic exploration of filesystem-based memory for LLM agents, formalizing roles of management, search, and execution agents around a shared memory store. It finds that organization primarily reduces retrieval cost but does not yet improve answer quality, and that tooling choices affect store shape as much as model selection.