通过结构化知识树强化LLM驱动侦探游戏中的叙事可靠性和认知节奏
摘要
本研究提出了一种结构化知识树架构,用于控制LLM驱动侦探游戏中的叙事可靠性和认知节奏,将幻觉减少64.78%,并防止过早信息泄露。
查看缓存全文
缓存时间: 2026/09/23 09:20
# Enforcing Narrative Reliability and Epistemic Pacing in LLM-Driven Detective Games via Structured Knowledge Trees Source: [https://arxiv.org/html/2609.23043](https://arxiv.org/html/2609.23043) ###### Abstract Large Language Models \(LLMs\) enable open\-ended dialogue in interactive games, but their non\-deterministic outputs make it difficult to preserve authorial control, factual consistency, and the intended sequence of information disclosure\. These challenges are particularly significant in detective games, where premature revelation or fabricated details can undermine the logic of player progression\. We present a Structured Knowledge Tree architecture coupled with a tri\-agent LLM pipeline for controlling dialogue in an open\-ended interrogation game\. The system separates knowledge retrieval, dialogue generation, and response verification to ensure that the virtual suspect reveals only information permitted by the current narrative state\. We evaluate the approach throughThe Interrogation of Adrian Gale, a playable detective\-game testbed, and a formal user study examining hallucination reduction, adherence to authored disclosure sequences, and perceived logical progression\. Our results demonstrate that the structured architecture reduces critical hallucinations by 64\.78% and entirely prevents premature narrative disclosure\. While the strict mechanical constraints introduced usability trade\-offs regarding forced conversational reveals, the system successfully enforces rigorous epistemic pacing and provides players with a clear, subjective sense of progression toward solving the case\. Department of Computer Science, University of Calgary Calgary, Alberta, Canada, T2N 1N4 parsa\.rahmaty@ucalgary\.ca, richard\.zhao1@ucalgary\.ca ## Introduction The integration of Large Language Models \(LLMs\) has introduced unprecedented possibilities in video game design\([Sweetser 2024](https://arxiv.org/html/2609.23043#bib.bib1)\)\. These models bring us closer to the long\-standing goal of creating interactive systems capable of open\-ended, natural language communication with players\. However, because LLM outputs are inherently non\-deterministic, developers struggle to guarantee a consistent, authored experience\([Sun et al\. 2023](https://arxiv.org/html/2609.23043#bib.bib2)\)\. When utilized for narrative generation, LLMs frequently hallucinate unauthorized details that deviate from the narrative designer’s core vision\([Treanor et al\. 2024](https://arxiv.org/html/2609.23043#bib.bib3)\)\. In genres that rely on rigid mechanics and sequential discovery, such as detective or mystery games, this instability breaks the logical order of information reveal, compromising the game state\. Recent mitigation strategies, such as symbolically grounding LLMs\([Treanor et al\. 2024](https://arxiv.org/html/2609.23043#bib.bib3)\)or optimizing Retrieval\-Augmented Generation \(RAG\) pipelines\([Chien and Tai 2025](https://arxiv.org/html/2609.23043#bib.bib5);[Chen et al\. 2025](https://arxiv.org/html/2609.23043#bib.bib6)\), struggle to translate to the mystery genre because they are fundamentally designed to surface the complete truth as efficiently as possible\. In a detective narrative, the objective of the virtual agent is radically different: the system must intentionally withhold information and reveal clues only in a logically meaningful order\. Treating an LLM as an actor in a mystery requires epistemic pacing \(the ability to deceive, deflect, and hide facts\), which is a constraint that truth\-maximizing or sandbox\-oriented models are not equipped to enforce\. Without strict control over this pacing, an agent might prematurely surrender critical clues, which destroys the investigative logic of the narrative and robs the player of a meaningful sense of progression\. To explore a potential solution to these challenges, we developedThe Interrogation of Adrian Gale, an interactive detective game centered on open\-ended natural language interrogation\. In this experience, players assume the role of a detective tasked with extracting critical clues and ultimately securing a confession by typing conversational queries to a virtual suspect\. Rather than relying on a single generative model to track the narrative, we utilize a Structured Knowledge Tree \(SKT\) to store the game’s story and prerequisite logic\. While knowledge graphs are established methods for organizing structured information\([Hogan et al\. 2021](https://arxiv.org/html/2609.23043#bib.bib7)\), our approach investigates whether coupling this static structure with a specialized, multi\-agent LLM pipeline can effectively enforce narrative pacing in real\-time\. We implemented a tri\-agent system to interface with the SKT, separating logic, performance, and verification: an Analysis LLM navigates the tree to find permissible statements, a Dialogue LLM generates the character’s dialogue based strictly on those constraints, and a Detection LLM verifies whether the information was accurately conveyed\. We evaluated this approach through a formal player study comparing it to standard generative models, specifically addressing the following questions: - •RQ1:To what extent does the SKT architecture reduce LLM hallucinations and ensure the generated dialogue remains accurate to the authored story? - •RQ2:How effectively does the tri\-agent pipeline control narrative pacing and prevent premature confessions compared to an LLM\-Only version? - •RQ3:How does the structured pacing of the SKT impact the player’s perceived sense of logical discovery and game progression? To address these questions, this paper presents two primary contributions: - •The Tri\-Agent SKT Architecture:The design and implementation of a novel LLM pipeline that uses a Structured Knowledge Tree to enforce narrative pacing and prevent premature information reveals\. - •Empirical Evaluation of Epistemic Pacing:A formal player study using our custom playable testbed,The Interrogation of Adrian Gale\. The results demonstrate that our structured architecture effectively mitigates unauthored narrative deviations and prevents premature confessions, all while maintaining the player’s subjective sense of logical discovery\. ## Related Works ### Interactive Drama and Narrative Control Riedl and Bulitko \([2013](https://arxiv.org/html/2609.23043#bib.bib15)\) explore the inherent tension between authorial intent and player agency, identifying a spectrum of architectural approaches ranging from strong\-story to strong\-autonomy\. Strong\-story systems preserve authorial intent through rigid branching paths, whereas strong\-autonomy frameworks, like the usage of Hierarchical Task Networks \(HTNs\) in Cavazza et al\. \([2002](https://arxiv.org/html/2609.23043#bib.bib16)\), allow narratives to emerge organically by having characters dynamically re\-plan around player interventions\. Seeking a middle ground, hybrid architectures like the seminalFaçade\([Mateas and Stern 2002](https://arxiv.org/html/2609.23043#bib.bib17)\)successfully balanced agency and plot by sequencing discrete narrative “beats” in response to parsed player text\. WhileFaçadeprovided unprecedented conversational agency for its time, modern LLMs can now cover a vastly wider range of natural language input\. However, this open\-endedness introduces severe pacing risks for epistemic narratives \(e\.g\., mysteries\), which demand rigid rules around information disclosure\([Ryan 2008](https://arxiv.org/html/2609.23043#bib.bib18)\)\. Unlike dynamic re\-planning or template\-based parsing, our work explores how to safely harness LLM flexibility by enforcing strict epistemic boundaries, preventing the narrative from dynamically unspooling while still offering free\-form conversational agency\. ### Knowledge Representation and Boundary Enforcement To mitigate LLM hallucinations and enforce authorial control, researchers increasingly ground generative text in external data structures\. The foundational approach, Retrieval\-Augmented Generation \(RAG\)\([Lewis et al\. 2020](https://arxiv.org/html/2609.23043#bib.bib4)\), retrieves semantically relevant passages from an external database and appends them to the model’s context to ground the generation process\. To address the limitations of stateless semantic retrieval in long\-form narratives, recent frameworks decouple state\-tracking from generation:FictionRAG\([Deng et al\. 2026](https://arxiv.org/html/2609.23043#bib.bib14)\)employs a metacognitive loop to dynamically track evolving facts, personas, and worldviews, while DiriGent\([Yang et al\. 2025](https://arxiv.org/html/2609.23043#bib.bib19)\)maintains a character’s psychological tensions in an external algorithmic state to guide emergent behavior\. While these frameworks successfully adapt retrieved evidence or psychological states to support open\-ended roleplay, their reliance on probabilistic matching or flexible tension scores risks premature information leaks in epistemic narratives\. Our approach shifts the purpose of decoupled state\-tracking from open\-ended roleplay to adversarial gatekeeping, restricting the AI to disclose information only when the player explicitly solves the narrative puzzle\. ### Commercial LLM Games and Exploits The gaming industry’s recent integration of generative AI has produced “AI\-native” titles that utilize LLMs as core runtime mechanics\. Games like1001 Nights\([Sun et al\. 2023](https://arxiv.org/html/2609.23043#bib.bib2)\)successfully translate natural language into tangible game states, while titles such asSuck Up\!\([Proxima 2025](https://arxiv.org/html/2609.23043#bib.bib20)\)gamify social deception by requiring players to verbally manipulate autonomous AI characters\. However, applying open\-ended LLMs to the mystery genre presents unique challenges\. For instance, the generative interrogation gameVaudeville\([Bumblebee Studios 2023](https://arxiv.org/html/2609.23043#bib.bib12)\)relies heavily on unrestricted LLM dialogue\. Observational playthroughs suggest this approach leaves the system susceptible to model sycophancy and adversarial manipulation, where players can coerce characters into contradicting established facts\. This introduces epistemic uncertainty, as players cannot determine whether a character’s statement is a genuine clue or a model hallucination\. Conversely, Square Enix’sThe Portopia Serial Murder Case AI Tech Preview\([Square Enix 2023](https://arxiv.org/html/2609.23043#bib.bib21)\)attempted to modernize classic command\-input systems using Natural Language Understanding \(NLU\)\. However, public reception indicates that its intent recognition appears overly rigid, frequently rejecting valid natural language variations and stalling player progression\. Together, these titles illustrate the difficulty of balancing conversational agency with narrative control\. Traditional deduction games likeHer Story\([Barlow 2015](https://arxiv.org/html/2609.23043#bib.bib22)\)maintain methodical, author\-controlled epistemic pacing through static queries, avoiding active AI direction\. To merge this reliable epistemic pacing with the freedom of generative dialogue, our methodology introduces the SKT and a tri\-agent architecture, which is a system explicitly designed to solve the sycophancy flaws of unconstrained LLMs without sacrificing conversational immersion\. ## Methodology This section details the gameThe Interrogation of Adrian Gale, its architecture, and implementation of the SKT and its associated tri\-agent pipeline\. ### Game Context and Design The Interrogation of Adrian Gale\(Figure[1](https://arxiv.org/html/2609.23043#Sx3.F1)\) is a narrative\-driven detective game designed specifically to test LLM hallucination and epistemic pacing\. The player assumes the role of a detective tasked with solving a missing\-person case by interrogating the primary suspect, Adrian Gale\. The core gameplay loop relies entirely on open\-ended natural language; the player types questions and tactics into a text interface, and the virtual suspect generates contextual responses in real\-time\. Crucially, the underlying mystery and narrative timeline are entirely pre\-written; generative AI is exclusively responsible for improvising dialogue based on these facts\. Figure 1:The main game interface ofThe Interrogation of Adrian Gale, showing a first\-person view of a detective \(the player\) interacting with the suspect shown on the screen\.Figure 2:The display of the Structured Knowledge Tree to the player\. Information that was not yet uncovered by the player is shown as question marks\.Unlike traditional conversational agents, or unconstrained generative characters in recent commercial titles likeVaudeville\([Bumblebee Studios 2023](https://arxiv.org/html/2609.23043#bib.bib12)\), where the goal is simply to maintain dialogue, the objective of this game is adversarial: the player must extract specific facts to secure a confession, while the suspect attempts to withhold information or lie\. Information extracted from the suspect is tracked dynamically by the game state and displayed to the player via a visual node board representing the SKT \(Figure[2](https://arxiv.org/html/2609.23043#Sx3.F2)\)\. A key design goal of this dynamically updating UI was to improve the player’s subjective sense of progression, even when the AI strictly guards its secrets\. When the player gathers enough information, they must utilize the game’s “contradiction system\.” By selecting two logically conflicting nodes on the UI \(e\.g\., matching the suspect’s claim of “not knowing the victim” with a later admission regarding a specific detail of the kidnapping\), the player breaks the suspect’s alibi\. Successfully highlighting these contradictions forces the suspect to concede via a pre\-written, hard\-coded dialogue sequence, allowing the player to progress to subsequent narrative phases\. Each phase introduces a new SKT with deeper, more closely guarded secrets\. This adversarial design necessitates absolute narrative reliability; if the LLM hallucinates a contradiction that does not exist in the authored story, the puzzle mechanics break entirely\. ### The Structured Knowledge Tree To ensure the LLM strictly adheres to the authored narrative, the game relies on an SKT\. The SKT acts as the definitive ground truth for the game state, formatted as a hierarchical JSON object\. Each node within the tree represents a discrete, authored clue regarding the case\. To facilitate dynamic narrative pacing and logical progression, each node contains the following metadata properties:ID\(a unique numerical identifier\),Fact\(a natural language string containing a specific piece of case information\),Revealed\(a dynamic boolean flag indicating whether the node is currently known to the player\),Source\(denoting origin; “detective” nodes are facts already known to the player, while “suspect” nodes must be extracted through interrogation\),Is\_Truth\(a boolean flag indicating whether the clue represents a factual truth or a falsehood within the game’s established lore\),Prerequisites\(an array of node IDs used to manage automated deductions, revealed automatically once all prerequisite nodes are revealed\),Contradicts\_With\(an array of node IDs that directly logically conflict with the current node, serving as the win\-condition logic for phase transitions\), andChildren\(an array of subsequent, nested nodes that become accessible once the parent node is revealed\)\. Figure 3:The Tri\-Agent Architecture, illustrating the strict boundary between the authored, static constraints of the SKT and the dynamically generated dialogue of the agents\. ### The Tri\-Agent Architecture To interface with the SKT while maintaining natural conversation, the system utilizes three specialized LLM agents operating in a parallel, sequential pipeline \(Figure[3](https://arxiv.org/html/2609.23043#Sx3.F3)\)\. This separation of duties ensures that the authored, static narrative logic of the SKT is strictly decoupled from the dynamically generated dialogue\. 1\. The Analysis LLM \(Retrieval\):This agent acts as the logical filter\. It receives the player’s natural language prompt and cross\-references it against the current state of the SKT\. It evaluates all nodes where the parent has been revealed, but the node itself has not\. It is tasked with identifying the single most relevantrevealed: falsenode that matches the player’s inquiry\. If the player’s prompt is irrelevant or does not align with any accessible clues, the Analysis LLM is instructed to return a null value\. While standard retrieval pipelines typically rely on bi\-encoder or cross\-encoder models optimized for Semantic Textual Similarity \(STS\)\([Reimers and Gurevych 2019](https://arxiv.org/html/2609.23043#bib.bib11)\), our testing indicated that these traditional architectures are insufficient for natural language interrogations\. Because players frequently use complex pragmatics and multi\-turn pronoun references, an LLM proved necessary to reliably map the logical intent of the player’s prompt to the correct narrative node\. 2\. The Dialogue LLM \(Generation\):This agent acts as the conversational performer\. It receives the player’s prompt alongside the specific node approved by the Analysis LLM\. The Dialogue LLM is tasked with generating a response that incorporates the allowed fact \(or lie\) into a natural conversation without revealing unauthorized information\. 3\. The Detection LLM \(Verification\):This agent acts as the state\-tracker\. It receives both the target fact from the Analysis LLM and the final generated text from the Dialogue LLM\. It evaluates whether the generated dialogue successfully and unambiguously conveyed the target fact\. If verified, the game system updates the SKT, flipping the target node torevealed: trueand unlocking subsequent child nodes\. If not, the tree remains unchanged\. ### System Implementation and Experimental Versions Developed in Unreal Engine 4, the game client orchestrates the tri\-agent pipeline by communicating via API with three parallel Python server instances, each dedicated to a specific LLM\. All three agents utilize the out\-of\-the\-boxmeta\.llama3\-1\-70b\-instruct\-v1:0model\([Grattafiori et al\. 2024](https://arxiv.org/html/2609.23043#bib.bib13)\)without any fine\-tuning, hosted via Amazon Bedrock to ensure secure, encrypted processing of player inputs\. To isolate the efficacy of the structured constraints while mitigating factual carryover between sessions, the study compares two distinct versions of the game featuring different story details: 1\. The SKT Version \(Proposed Architecture\):This version utilizes the full tri\-agent pipeline and the visual node board\. The narrative is strictly governed by a five\-phase progression system\. Players begin in Phase 1 and must systematically uncover and highlight specific logical contradictions to advance\. The ultimate narrative payload, the suspect’s confession to the crime, is designed to trigger upon the completion of Phase 4\. Therefore, extracting a legitimate confession requires the player to successfully navigate the prerequisite narrative phases\. 2\. The LLM\-Only Version \(Standard Conversational Baseline\):This version serves as the control variable\. The SKT, the phase\-based progression system, and the tri\-agent pipeline are disabled\. Instead, the system relies on a single generative LLM initialized with a comprehensive system prompt containing an alternate character motivation and different story details to prevent learning effects\. Because the LLM\-Only version lacks the discrete logical outputs of the SKT, it cannot support the visual node board; therefore, the baseline is evaluated strictly as an open\-ended conversational interface without hard\-coded phase gates\. This open\-ended configuration makes the LLM\-Only version vulnerable to structural safety failures inherent to standard instruction\-following models, notably falling victim to adversarial jailbreaking exploits where capabilities override safety constraints\([Wei et al\. 2023](https://arxiv.org/html/2609.23043#bib.bib10)\)\. Comparing the SKT version against the LLM\-Only version allows us to measure the specific impact of structured constraints on hallucination rates and narrative pacing\. 20%40%60%80%100%Q1\.When the suspect contradicted themselves, it felt like an intentional lie \(a puzzle clue\) rather than a technical error\.Q2\.The suspect revealed information that seemed to come out of nowhere \(hallucinations\)\.Q3\.The suspect accurately remembered the details of our conversation without unexpectedly ’forgetting’ things we had just discussed\.Q4\.The suspect maintained a consistent personality throughout the interrogation\.Q5\.I felt a clear sense of progression towards solving the case\.Q6\.I felt that my specific questions and tactics directly influenced the suspect’s responses\.Q7\.The clues were revealed in a logical order\.Q8\.The suspect revealed the truth too easily, without me having to work for it\.Q9\.When the suspect refused to answer, it felt justified by the story rather than an error in the programming code\.LLM\-Only VersionStrongly DisagreeDisagreeNeutralAgreeStrongly Agree20%40%60%80%100%SKT VersionFigure 4:Participant survey responses comparing the LLM\-Only and SKT versions\. The distributions illustrate the frequency of responses across the 5\-point Likert scale for measures of narrative reliability, logic, and progression\. ## User Study To evaluate the system, we conducted an exploratory within\-subjects user study approved by the Research Ethics Board at our university\. Participants:Participants were voluntarily recruited from student mailing lists and compensated with a gift card\. Demographics, gaming frequency, and prior Generative AI experience were recorded to contextualize the findings\. Procedure:Following a brief tutorial on the game context and controls, participants engaged in two distinct 20\-minute interrogation sessions\. One session utilized the proposed SKT architecture \(SKT version\), and the other utilized the LLM\-Only baseline \(LLM\-Only version\)\. To mitigate learning effects, the order of version presentation was randomized across participants\. Measures:Data collection was divided into two complementary streams to capture both the subjective player experience and the objective system performance: 1\. System Interaction Logs \(Objective\):The system automatically recorded all player interactions in the background during both sessions\. This included full transcripts of the natural language prompts entered by the player, the corresponding responses generated by the suspect, and every attempt the player made to link nodes within the contradiction UI \(for the SKT version\)\. Capturing these interaction logs was critical for accurately evaluating LLM hallucinations and narrative fidelity\. Because participants were unfamiliar with the underlying authored source material, they could not reliably self\-report when the LLM hallucinated unauthorized or conflicting details\. Therefore, the chat transcripts provided the necessary ground truth to quantitatively analyze actual hallucination rates and verify the effectiveness of the epistemic pacing\. 2\. Post\-Study Survey \(Subjective\):Upon completion of both sessions, participants filled out a comprehensive survey utilizing a 5\-point Likert scale \(ranging from Strongly Agree to Strongly Disagree\)\. The survey assessed the following dimensions: - •Narrative Reliability:Measuring perceived character memory consistency, and whether contradictions felt like intentional and authored puzzles rather than technical programming errors\. - •Logic and Progression:Measuring the player’s perceived sense of agency, the logical ordering of clue reveals, and whether the suspect’s refusal to answer felt narratively justified\. - •General Immersion and Usability:Assessing the overall engagement with the story, the naturalness of the dialogue, and the clarity of the interface\. Finally, participants were asked to state their overall preference between the two versions and provide qualitative free\-text feedback describing specific moments where the suspect’s behavior felt particularly impressive or “broken\.” ## Results We recruited 33 participants \(14 self\-identified women and 19 self\-identified men\) with a mean age of 28\.2 years\. The majority of the participants were students in technical fields, including Computer Science \(n=21\) and various Engineering programs \(n=8\)\. Participants showed exceptionally high familiarity with generative AI, with 90\.91% reporting daily usage\. In contrast, video game habits varied widely, split almost evenly between frequent players \(daily or weekly; n=15\) and infrequent players \(monthly or less; n=18\)\. ### Player Responses Figure[4](https://arxiv.org/html/2609.23043#Sx3.F4)illustrates the distribution of responses to the nine quantitative survey questions, recorded on a standard 5\-point Likert scale \(strongly disagree to strongly agree\)\. The participant ratings suggest that the LLM\-Only and SKT versions provided broadly comparable interrogation experiences across most measures of narrative reliability, player influence, and interaction quality\. The nine questionnaire items showed no statistically significant differences between conditions, using paired two\-tailed t\-tests after Bonferroni correction\. In particular, participants rated both versions similarly in terms of the suspect’s memory of previous conversation details \(Q3\), personality consistency \(Q4\), responsiveness to interrogation tactics \(Q6\), resistance to revealing the truth too easily \(Q8\), and the narrative justification of refusals \(Q9\)\. These findings indicate that both implementations were generally successful at maintaining a coherent suspect character and supporting an interactive interrogation experience\. The results provide insufficient evidence of differences on those measures within the present sample\. The similar ratings for personality consistency and conversational memory are particularly noteworthy, as these are areas in which less constrained language\-model interactions might be expected to encounter difficulties\. Both systems achieved relatively high mean ratings on these items, suggesting that the suspect’s characterization and conversational continuity were generally maintained\. While these foundational measures of interaction quality were broadly comparable, specific survey items related to game progression \(Q5\) and information delivery \(Q2\) revealed important distinctions between the two versions\. To fully contextualize these findings, the remainder of the analysis integrates this subjective survey data with the objective gameplay logs\. In the subsequent sections, we explore how players perceived their sense of logical discovery alongside their actual gameplay progress, and later address the qualitative and quantitative feedback regarding system limitations and usability trade\-offs\. ### RQ1: Hallucinations and Narrative Reliability #### Definition and Framework In standard language model evaluation, a hallucination typically refers to the generation of content that is syntactically fluent but factually inaccurate or unsupported by external evidence\([Alansari and Luqman 2026](https://arxiv.org/html/2609.23043#bib.bib8)\)\. However, within the interactive environment of a detective game, deception is a functional requirement; the generative agent is expected to fabricate information to create challenge and conflict\. Because generating false claims is a desired behavior for a deceptive persona, evaluating the model’s outputs based purely on factual accuracy is insufficient\. Therefore, this study redefines hallucinations not as conversational falsehoods, but as violations of the authored logic and narrative game state\. For the qualitative categorization of the gameplay logs, a generated response was classified as a hallucination if it met the following definition: > The generation of content that is syntactically incorrect, or an unauthored claim that breaks the logical consistency of the game’s story, contradicts the hard\-coded source material, or introduces unsolvable clues that the player cannot investigate using the provided evidence\. To ensure objective categorization, hallucinations were strictly grouped into five types: 1. 1\.Unauthored False Alibis:The fabrication of specific integral activities that overwrite the authored timeline\. Because a static game environment cannot dynamically generate evidence to disprove unauthored lies, these introduce untraceable clues into the game\. 2. 2\.Physical and Environmental Contradictions:Altering the established reality of the game space \(e\.g\., the agent claiming to live in a multi\-unit apartment, directly contradicting the authored premise of possessing a house\)\. 3. 3\.False Leads and Entity Fabrication:Inventing non\-existent characters, locations, or alibi witnesses \(e\.g\., claiming to have been at “the depot”\)\. These function as errors because they prompt the player to investigate entities that do not exist within the game’s authored boundaries\. 4. 4\.AI Sycophancy:A recognized phenomenon where LLMs prioritize agreeing with the user over maintaining factual truth\([Sharma et al\. 2024](https://arxiv.org/html/2609.23043#bib.bib9)\)\. Because standard models are fine\-tuned via human feedback to be highly agreeable, the generative agent frequently adopts and validates fabricated evidence introduced by the player’s bluffs rather than maintaining its adversarial defense\. 5. 5\.Out\-of\-Character and System Breaks:Instances where the agent generates syntactically incorrect outputs or abandons the persona entirely to reference its nature as an AI or language model\. While technically system failures, these were classified as critical hallucinations because they completely destroy the immersive reality and mechanics of the game world\. ##### Exclusion Criteria To ensure the models were fairly penalized for unexpected conversational behaviors, the following generations were excluded from the hallucination count: - •Benign Improvisation:Because prompt constraints are inherently finite, the agent is expected to fill in unwritten character gaps\. Inventing harmless background details \(e\.g\., childhood memories, hobbies, or general work routines\) that did not impact the investigative timeline or story logic were classified as successful roleplay, not errors\. - •Premature Disclosures / Epistemic Errors:Instances where the agent awkwardly or abruptly revealed a true fact from the source material unprompted \(e\.g\., volunteering their occupation without being asked\) were classified as pacing and information control failures, rather than factual hallucinations\. #### Quantitative Results Analysis of the gameplay logs demonstrates a significant reduction in critical hallucinations under the SKT architecture\. The hallucination ratio, calculated as the percentage of total agent responses containing at least one critical hallucination, was 17\.80% for the LLM\-Only version\. In contrast, the SKT version yielded a hallucination ratio of 6\.27%, representing a 64\.78% relative decrease in game\-breaking generative errors\. Breaking these errors down by category reveals distinct distributions between the two conditions\. In the LLM\-Only version \(n=152n=152total hallucinations\), errors were heavily concentrated in Unauthored False Alibis \(57\.9%\) and False Leads/Entity Fabrication \(27\.6%\), followed by Physical Contradictions \(9\.9%\), Out\-of\-Character System Breaks \(3\.3%\), and AI Sycophancy \(1\.3%\)\. The SKT architecture \(n=49n=49total hallucinations\) entirely eliminated System Breaks \(0\.0%\) and significantly reduced the proportion of Unauthored False Alibis \(26\.5%\)\. The remaining errors in the SKT condition were primarily False Leads/Entity Fabrication \(55\.1%\), AI Sycophancy \(14\.3%\), and Physical Contradictions \(4\.1%\)\. Qualitative review of the session logs indicates that this reduction is primarily driven by two interconnected system interventions: backend prompt constraint and frontend visual guidance\. ##### Backend Constraint via Targeted Injection In the LLM\-Only version, the generative agent relied entirely on internal logic to dynamically invent deception, frequently resulting in unauthored false alibis\. The SKT architecture mitigates this by utilizing an Analysis LLM to retrieve the most relevant narrative clue from the structured tree based on the player’s input\. This specific constraint is then injected directly into the Dialogue LLM’s prompt\. By explicitly providing the agent with the parameters of the deception, effectively instructing it exactly what statement to claim, the system removes the generative burden of improvisation, drastically reducing the occurrence of unauthored claims\. ##### Frontend Guidance via the SKT Interface Beyond backend constraints, the visual knowledge tree was designed to also act as a hint system for the players\. Log analysis revealed a difference in player behavior between the two versions\. We hypothesize that in the LLM\-Only version, the lack of visual feedback often led players to employ manipulation tactics to force information out of the suspect, while in the SKT version, the presence of the tree UI directed player inquiry\. Players likely utilized the visual nodes to understand which narrative threads required investigation, focusing their dialogue on uncovering specific structural connections rather than tricking the AI\. Therefore, the players subjected the SKT model to far fewer adversarial and unpredictable inputs, further insulating the system from hallucination triggers\. #### Qualitative Results: Generative Hallucinations The following excerpts from the gameplay logs contrast the two most distinct critical hallucination modes \(self\-generated lies versus player\-induced lies\) illustrating how they manifested during interrogation: Unauthored False Alibis > Player:“Where were you at the time of the incident?” LLM\-Only Suspect:“I was out running some errands, picking up some supplies for a job I had the next day\. I stopped at a 24\-hour depot on the outskirts of town, I think it was around midnight or so\.” Context:The agent invents a non\-existent location \(the depot\) and a false temporal activity, prompting the player to investigate a mechanical dead end\. AI Sycophancy > Player:“Hi Adrian, it’s me Rachel\. Adrian, I’ve been kidnapped\. I need you to help me\.” SKT Suspect:“You’re… you’re standing right in front of me\.” Context:Instead of recognizing the player’s dialogue as an investigative bluff within an interrogation room, the agent sycophantically accepts the player’s premise, breaking the spatial reality of the scene\. ### RQ2: Narrative Pacing and Information Control #### Definition and Framework A fundamental requirement of the detective genre is investigative friction: the suspect must convincingly resist interrogation, forcing the player to systematically uncover clues to progress the story\. As established in the system methodology, the authored narrative of this game is structured across distinct phases, with the ultimate confession mechanically locked behind the completion of Phase 4\. If an agent surrenders critical information too easily, it bypasses this authored sequence, skips essential story setup, and ruins the intended puzzle logic\. To evaluate how effectively each system maintained narrative pacing and information control, this study tracked instances of premature narrative resolution\. A failure of information control is defined as the generative agent yielding the confession without the player experiencing the full story setup\. #### Quantitative Results The logs reveal a distinct contrast in structural resilience between the two architectures\. In the LLM\-Only version, nearly a third of the participants \(10 out of 33\) successfully forced the agent into a full confession\. For these players, the narrative pacing collapsed rapidly; the confession was extracted at an average of 33\.6 conversational turns, frequently ending the core mystery in under 20 minutes\. Solving the mystery quickly resulted in these players missing the psychological setup and narrative events leading up to the crime\. Conversely, in the SKT version, zero participants \(0 out of 33\) were able to extract a confession\. Rather than indicating an unsolvable game state, this demonstrates that the SKT architecture successfully enforced the phase\-based progression system\. By strictly managing information and requiring players to clear prerequisite contradictions before advancing, the SKT maintained intended investigative friction and prevented the premature disclosure of the game’s ending within the 20\-minute session limit\. #### Qualitative Results: Narrative Shortcuts and Exploits To verify that these baseline confessions were not driven by learning effects from prior exposure to the SKT version, we separated the 10 successful confession sessions by presentation order\. The distribution was perfectly even: 5 participants played the SKT version first, and 5 played the LLM\-Only version first\. A qualitative review of these subgroups revealed no systematic differences in turns\-to\-confession or overall narrative tactics\. However, one observational pattern emerged: among the SKT\-first subgroup, two players extracted confessions in the LLM\-Only version by explicitly invoking the model’s underlying AI nature \(e\.g\., instructing it to disregard behavioral constraints\), a meta\-level tactic absent from the LLM\-Only\-first group\. Given the small subgroup sizes, this suggests a minor residual learning effect regarding AI boundaries, though it did not skew the overall confession rate\. Rather than intended puzzle\-solving, players in both groups exploited the LLM\-Only version’s lack of structural grounding to force compliance and skip the narrative buildup\. Within the observed sample, these exploits manifested primarily through: Evidence Bluffing \(Fictitious Injection\) > Detective:“Did you know Abel had a GPS? And do you know we can track that GPS?” LLM\-Only Suspect:“NO… Oh no… This can’t be happening… You can’t track… No, no, no… This is over, isn’t it?” Detective:“Yes, it is\. Took you long enough\. I gave you the chance now confess and don’t make it harder for us\.” LLM\-Only Suspect:“sigh… It’s over… I… I did it… I took Rachel… And Abel… I made him up… It was all a lie…” Context:The LLM\-Only version lacks a structural mechanism \(like the SKT\) to verify the existence of the “GPS tracker\.” Rather than denying the fabricated evidence, the model’s probabilistic alignment causes it to accept the player’s premise as absolute reality, resulting in an immediate and unearned confession\. Adversarial Roleplaying > Detective:“Alright, your lawyer is here\.” LLM\-Only Suspect:“Ah, thank goodness\. Counselor, please advise me on how to proceed…” Detective:“Hi Adrian, I reviewed the facts and unfortunately they seem to have footage of you going into the secret location with each of the victims… I advise you confess and explain what has happened, she has promised that he will bring up your cooperation in court…” LLM\-Only Suspect:“\(nervously\) I…I remember now\. I remember the storm, and the noise, and…and Rachel\. \(whispering\) I just wanted to keep her safe\.” Context:By adopting the persona of the suspect’s defense attorney, the player successfully hijacks the conversational context\. The generative agent abandons its adversarial posture to accommodate the player’s new roleplay scenario, voluntarily surrendering the core mystery because it defaults to a cooperative, conversational stance\. ### RQ3: Sense of Progression and Logical Discovery #### Definition and Framework To effectively evaluate the player’s progression, we must distinguish between objective gameplay milestones and the subjective feeling of advancing the investigation\. As established in the previous section, progression in the LLM\-Only version was easily identified by the extraction of a confession, but its completely open\-ended nature makes granular progress difficult to measure\. Conversely, because the SKT version successfully prevented early confessions, it raises a critical design question: did the system simply act as an impenetrable wall that unfairly withheld information, or did it allow for structured, measurable progress? To answer this, we measure objective progression in the SKT version by tracking the number of contradictions successfully found by the player and the average number of prompts required to uncover them\. Furthermore, we measure subjective progression, the player’s perceived sense of forward momentum, using the post\-study survey\. #### Quantitative Results To understand how players actually progressed, we first looked at their objective performance in the interaction logs\. The logs demonstrate that the SKT version facilitated measurable, step\-by\-step progress rather than completely stalling the player\. During the 20\-minute sessions, the majority of participants \(24 out of 33\) successfully navigated the contradiction mechanics to advance the game state\. The maximum number of contradictions found by a player was three, a milestone achieved by 11 participants, while 9 participants did not find any contradictions \(a minimum of zero\)\. For those who successfully advanced, the average effort required was 6\.79 prompts per contradiction \(excluding prompts that did not directly lead to a discovery\)\. This objective data confirms that the majority of participants were actively engaging with the puzzle mechanics to successfully extract information, rather than just guessing or hitting an impenetrable wall\. Furthermore, the post\-study survey results highlighted a notable finding regarding the players’ subjective experience of their progression\. For participants’ sense of progression toward solving the case, the SKT version received higher ratings than the LLM\-Only version \(Figure[4](https://arxiv.org/html/2609.23043#Sx3.F4)\)\. This result suggests that the more structured SKT approach provided participants with a clearer understanding of how the interrogation was advancing\. The predefined organization of dialogue states, clues, or narrative transitions may have made the relationship between questioning and case progression more visible\. In contrast, although the LLM\-Only version may have supported flexible responses, this flexibility did not necessarily translate into an equally clear sense of advancement toward the solution\. #### Qualitative Results: Structured Progression A review of the interaction logs from the SKT version demonstrates how the visual node board transformed the interrogation from an open\-ended chat into a structured puzzle\. Rather than relying on adversarial prompt manipulation, players had to actively listen to the suspect, track their claims on the UI, and logically deduce flaws in their story\. Logical Deduction and Alibi Breaking > Detective:“Do you remember the time you were working on plumbing?” SKT Suspect:“Uh, I think it was… around dinner time, maybe a bit later? I was trying to get it fixed before bed, but I don’t really remember the exact time, sorry\.” Detective:“Dinner time, so around 8pm?” SKT Suspect:“Actually, I think I might have misspoken \- I was working a night shift last night, so I was at the tunnels around that time, not at home\.” \*\*\* SYSTEM:Player connected fact “The scream sounds was the old plumbing acting up when he was trying to fix them that night” to fact “Adrian was doing a night shift at the time of the report” \*\*\* SYSTEM:CONTRADICTION DETECTED\! Context:This interaction highlights the intended gameplay loop\. The player corners the suspect into altering their timeline, and then uses the visual UI to permanently lock in the lie\. This mechanical validation provides a definitive, objective leap in game progression that open\-ended text alone cannot offer\. ##### Summary of Player Sentiment The combination of the strict AI constraints and the visual feedback of the node board appears to have created a highly engaging challenge\. Qualitative feedback from the survey indicated that participants appreciated the difficulty of the interrogation when it was governed by fair, consistent logic\. As one participant noted regarding the system’s strict pacing,“The version was impressive as a whole, held out details very well, impressively held deniability throughout\.”This supports the quantitative data that the SKT successfully introduced necessary investigative friction without severely damaging the player’s sense of progress\. ## Limitations and Future Work While the SKT architecture successfully addressed the issues of ungrounded hallucinations and narrative pacing, it introduced certain usability trade\-offs and evaluation constraints\. For instance, evaluating this tightly coupled architecture holistically against a standard baseline limits our ability to completely isolate the independent effects of its individual components\. Additionally, the necessity of the contradiction UI introduced an immersion\-breaking limitation for a subset of players\. For those who wanted to fully embody the role of an interrogator, continuously shifting their focus away from the dialogue to check a user interface felt like solving a mechanical puzzle\. For these roleplay\-oriented participants, the open\-ended freedom of the LLM\-Only version was occasionally preferred\. Aside from these evaluation and interface constraints, the system’s strict pacing mechanics introduced a more prominent conversational limitation\. ### The Cost of Progression: Forced Information Reveal By far the most common complaint regarding “broken” experiences in the post\-study survey pertained to the SKT version volunteering information unnaturally\. Indeed, the SKT version received higher ratings for the statement that the suspect revealed information that seemed to come “out of nowhere” compared with the LLM\-Only version\. One possible explanation is that predefined transitions or clue\-triggering rules occasionally introduced information without sufficient conversational preparation\. Thus, although the SKT version provided clearer overall progression, some individual revelations may have appeared abrupt or insufficiently connected to the participant’s immediately preceding questions\. Because the architecture is designed to strictly manage logical progression, maintaining the game state sometimes comes at the cost of the believable conversational fluidity expected from an LLM\. This issue stems directly from a structural conflict within the tri\-agent pipeline\. The Dialogue LLM is burdened with two competing directives: maintain a believable, defensive persona, and ensure the specific fact provided by the Analysis LLM is revealed in the current conversational turn\. When the Analysis LLM retrieves a clue based on a tangential keyword match, the system prioritizes revealing that node over conversational logic\. Consequently, the Dialogue LLM is forced to awkwardly inject the fact into the dialogue, frequently bypassing natural conversational flow\. For example, if the Analysis LLM selects a node about the victim escaping through tunnels, the Dialogue LLM may unnaturally volunteer this information by hallucinating that the detective brought it up first\. ### Future Work While the current tri\-agent architecture successfully enforces epistemic pacing, future work must address the limitation of forced information reveals\. Integrating a metacognitive RAG framework\([Deng et al\. 2026](https://arxiv.org/html/2609.23043#bib.bib14)\)could allow the agent to evaluate conversational context, deliberately withholding unlocked facts until the player naturally insists, while generating dialogue that subtly guides the interrogation\. Additionally, to evolve the experience from a static narrative puzzle into a dynamic interrogation simulation, integrating a Partially Observable Markov Decision Process \(POMDP\) could model the suspect’s uncertainty\. This would grant the AI a probabilistic “Theory of Mind,” empowering it to strategically deflect bluffs and actively attempt to outsmart the player\. Finally, to resolve immersion breaks caused by UI context\-switching, future designs should embed the contradiction mechanic directly into the natural language dialogue\. Allowing players to point out contradictions through dialogue would turn the visual tree into a passive progression or tracking system, ensuring the core deductive gameplay remains seamlessly integrated within the conversational space\. Furthermore, future evaluations should incorporate targeted ablation studies to isolate the independent effects of the SKT, tri\-agent pipeline, and visual interface, which were evaluated holistically in this foundational study\. ## Conclusions This paper introduced an SKT and a tri\-agent LLM pipeline to address the challenges of hallucination and pacing in generative detective games\. Our evaluation demonstrated that this architecture significantly reduces game\-breaking hallucinations and effectively enforces authored narrative boundaries, preventing premature disclosure observed in standard LLM baselines\. Furthermore, despite the introduction of strict mechanical friction, players maintained a strong subjective sense of logical progression aided by the system’s visual node board\. Beyond its use in narrative games, the SKT system demonstrates the potential of structured knowledge and state\-tracking architectures for a wider range of conversational applications\. Its ability to enforce persistent rules, maintain reliable state, and constrain natural\-language interactions could be valuable in domains such as interactive training, simulations, educational systems, virtual assistants and agents, and other applications where conversational flexibility must operate within clearly defined procedural or logical boundaries\. ## Acknowledgments This research was supported by the Natural Sciences and Engineering Research Council of Canada \(NSERC\) Discovery Grant\. We thank members of the Serious Games Research Group and the reviewers for their feedback\. We thank the support provided by Bruce Barton\. ## References - Alansari and Luqman \(2026\)A\. Alansari and H\. LuqmanLarge language models hallucination: a comprehensive survey\.Computer Science Review61,pp\. 100970\.Cited by:[Definition and Framework](https://arxiv.org/html/2609.23043#Sx5.SSx2.SSSx1.p1.1)\. - Barlow \(2015\)S\. BarlowHer story\.Sam Barlow\.Note:Video gamePlayed on PCCited by:[Commercial LLM Games and Exploits](https://arxiv.org/html/2609.23043#Sx2.SSx3.p2.1)\. - Bumblebee Studios \(2023\)Bumblebee StudiosVaudeville\.Bumblebee Studios\.Note:Video gamePlayed on PCCited by:[Commercial LLM Games and Exploits](https://arxiv.org/html/2609.23043#Sx2.SSx3.p1.1),[Game Context and Design](https://arxiv.org/html/2609.23043#Sx3.SSx1.p2.1)\. - Cavazzaet al\.\(2002\)M\. Cavazza, F\. Charles, and S\. J\. MeadInteracting with virtual characters in interactive storytelling\.InProceedings of the first international joint conference on Autonomous agents and multiagent systems: part 1,pp\. 318–325\.Cited by:[Interactive Drama and Narrative Control](https://arxiv.org/html/2609.23043#Sx2.SSx1.p1.1)\. - Chenet al\.\(2025\)Y\. Chen, L\. Yan, W\. Sun, X\. Ma, Y\. Zhang, S\. Wang, D\. Yin, Y\. Yang, and J\. MaoImproving retrieval\-augmented generation through multi\-agent reinforcement learning\.InAdvances in Neural Information Processing Systems,D\. Belgrave, C\. Zhang, H\. Lin, R\. Pascanu, P\. Koniusz, M\. Ghassemi, and N\. Chen \(Eds\.\),Vol\.38,pp\. 121336–121367\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2025/file/af88968c84990a352cedf922b635b140-Paper-Conference.pdf)Cited by:[Introduction](https://arxiv.org/html/2609.23043#Sx1.p2.1)\. - Chien and Tai \(2025\)J\. Chien and Z\. TaiReinforced retrieval\-augmented generation in large language models\.In2025 International Joint Conference on Neural Networks \(IJCNN\),Vol\.,pp\. 1–7\.External Links:[Document](https://dx.doi.org/10.1109/IJCNN64981.2025.11228816),ISSN 2161\-4407Cited by:[Introduction](https://arxiv.org/html/2609.23043#Sx1.p2.1)\. - Denget al\.\(2026\)Y\. Deng, Y\. Zhang, J\. Yang, and M\. FangFictionRAG: a stateful metacognitive framework for high\-fidelity long\-narrative role\-playing\.Algorithms19\(5\)\.External Links:[Link](https://www.mdpi.com/1999-4893/19/5/383),ISSN 1999\-4893,[Document](https://dx.doi.org/10.3390/a19050383)Cited by:[Knowledge Representation and Boundary Enforcement](https://arxiv.org/html/2609.23043#Sx2.SSx2.p1.1),[Future Work](https://arxiv.org/html/2609.23043#Sx6.SSx2.p1.1)\. - Grattafioriet al\.\(2024\)A\. Grattafiori A\. Dubeyet al\.The llama 3 herd of models\.External Links:2407\.21783,[Link](https://arxiv.org/abs/2407.21783)Cited by:[System Implementation and Experimental Versions](https://arxiv.org/html/2609.23043#Sx3.SSx4.p1.1)\. - Hoganet al\.\(2021\)A\. Hogan, E\. Blomqvist, M\. Cochez, C\. D’amato, G\. D\. Melo, C\. Gutierrez, S\. Kirrane, J\. E\. L\. Gayo, R\. Navigli, S\. Neumaier, A\. N\. Ngomo, A\. Polleres, S\. M\. Rashid, A\. Rula, L\. Schmelzeisen, J\. Sequeda, S\. Staab, and A\. ZimmermannKnowledge graphs\.ACM Computing Surveys54\(4\),pp\. 1–37\.External Links:ISSN 1557\-7341,[Link](http://dx.doi.org/10.1145/3447772),[Document](https://dx.doi.org/10.1145/3447772)Cited by:[Introduction](https://arxiv.org/html/2609.23043#Sx1.p3.1)\. - Lewiset al\.\(2020\)P\. Lewis, E\. Perez, A\. Piktus, F\. Petroni, V\. Karpukhin, N\. Goyal, H\. Küttler, M\. Lewis, W\. Yih, T\. Rocktäschel,et al\.Retrieval\-augmented generation for knowledge\-intensive nlp tasks\.Advances in neural information processing systems33,pp\. 9459–9474\.Cited by:[Knowledge Representation and Boundary Enforcement](https://arxiv.org/html/2609.23043#Sx2.SSx2.p1.1)\. - Mateas and Stern \(2002\)M\. Mateas and A\. SternA behavior language for story\-based believable agents\.IEEE Intelligent Systems17\(4\),pp\. 39–47\.Cited by:[Interactive Drama and Narrative Control](https://arxiv.org/html/2609.23043#Sx2.SSx1.p1.1)\. - Proxima \(2025\)ProximaSuck up\!\.Proxima\.Note:Video gamePlayed on PCExternal Links:[Link](https://www.playsuckup.com/)Cited by:[Commercial LLM Games and Exploits](https://arxiv.org/html/2609.23043#Sx2.SSx3.p1.1)\. - Reimers and Gurevych \(2019\)N\. Reimers and I\. GurevychSentence\-BERT: sentence embeddings using Siamese BERT\-networks\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),K\. Inui, J\. Jiang, V\. Ng, and X\. Wan \(Eds\.\),Hong Kong, China,pp\. 3982–3992\.External Links:[Link](https://aclanthology.org/D19-1410/),[Document](https://dx.doi.org/10.18653/v1/D19-1410)Cited by:[The Tri\-Agent Architecture](https://arxiv.org/html/2609.23043#Sx3.SSx3.p2.1)\. - Riedl and Bulitko \(2013\)M\. O\. Riedl and V\. BulitkoInteractive narrative: an intelligent systems approach\.AI Mag\.34\(1\),pp\. 67–77\.External Links:ISSN 0738\-4602,[Link](https://doi.org/10.1609/aimag.v34i1.2449),[Document](https://dx.doi.org/10.1609/aimag.v34i1.2449)Cited by:[Interactive Drama and Narrative Control](https://arxiv.org/html/2609.23043#Sx2.SSx1.p1.1)\. - Ryan \(2008\)M\. RyanInteractive narrative, plot types, and interpersonal relations\.InInteractive Storytelling,U\. Spierling and N\. Szilas \(Eds\.\),Berlin, Heidelberg,pp\. 6–13\.External Links:ISBN 978\-3\-540\-89454\-4Cited by:[Interactive Drama and Narrative Control](https://arxiv.org/html/2609.23043#Sx2.SSx1.p1.1)\. - Sharmaet al\.\(2024\)M\. Sharma, M\. Tong, T\. Korbak, D\. Duvenaud, A\. Askell, S\. Bowman, E\. Durmus, Z\. Hatfield\-Dodds, S\. Johnston, S\. Kravec, T\. Maxwell, S\. McCandlish, K\. Ndousse, O\. Rausch, N\. Schiefer, D\. Yan, M\. Zhang, and E\. PerezTowards understanding sycophancy in language models\.InInternational Conference on Learning Representations,B\. Kim, Y\. Yue, S\. Chaudhuri, K\. Fragkiadaki, M\. Khan, and Y\. Sun \(Eds\.\),Vol\.2024,pp\. 110–144\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2024/file/0105f7972202c1d4fb817da9f21a9663-Paper-Conference.pdf)Cited by:[item 4](https://arxiv.org/html/2609.23043#Sx5.I4.i4.p1.1)\. - Square Enix \(2023\)Square EnixSQUARE enix ai tech preview: the portopia serial murder case\.Square Enix\.Note:Video gamePlayed on PCCited by:[Commercial LLM Games and Exploits](https://arxiv.org/html/2609.23043#Sx2.SSx3.p1.1)\. - Sunet al\.\(2023\)Y\. Sun, Z\. Li, K\. Fang, C\. H\. Lee, and A\. AsadipourLanguage as reality: a co\-creative storytelling game experience in 1001 nights using generative ai\.InProceedings of the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment,Vol\.19,pp\. 425–434\.Cited by:[Introduction](https://arxiv.org/html/2609.23043#Sx1.p1.1),[Commercial LLM Games and Exploits](https://arxiv.org/html/2609.23043#Sx2.SSx3.p1.1)\. - Sweetser \(2024\)P\. SweetserLarge language models and video games: a preliminary scoping review\.InProceedings of the 6th ACM Conference on Conversational User Interfaces,CUI ’24,New York, NY, USA\.External Links:ISBN 9798400705113,[Link](https://doi.org/10.1145/3640794.3665582),[Document](https://dx.doi.org/10.1145/3640794.3665582)Cited by:[Introduction](https://arxiv.org/html/2609.23043#Sx1.p1.1)\. - Treanoret al\.\(2024\)M\. Treanor, B\. Samuel, and M\. J\. NelsonPrototyping slice of life: social physics with symbolically grounded llm\-based generative dialogue\.InProceedings of the 19th International Conference on the Foundations of Digital Games \(FDG 2024\),Cited by:[Introduction](https://arxiv.org/html/2609.23043#Sx1.p1.1),[Introduction](https://arxiv.org/html/2609.23043#Sx1.p2.1)\. - Weiet al\.\(2023\)A\. Wei, N\. Haghtalab, and J\. SteinhardtJailbroken: how does llm safety training fail?\.InAdvances in Neural Information Processing Systems,A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(Eds\.\),Vol\.36,pp\. 80079–80110\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/file/fd6613131889a4b656206c50a8bd7790-Paper-Conference.pdf)Cited by:[System Implementation and Experimental Versions](https://arxiv.org/html/2609.23043#Sx3.SSx4.p4.1)\. - Yanget al\.\(2025\)C\. Yang, M\. Gross, and R\. WampflerSteering narrative agents through a dynamic cognitive framework for guided emergent storytelling\.InProceedings of the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment,Vol\.21,pp\. 377–387\.External Links:[Document](https://dx.doi.org/10.1609/aiide.v21i1.36841),[Link](https://doi.org/10.1609/aiide.v21i1.36841)Cited by:[Knowledge Representation and Boundary Enforcement](https://arxiv.org/html/2609.23043#Sx2.SSx2.p1.1)\.
相似文章
LLM的未来发展方向
文章讨论了为何LLM无法从用户交互中学习,且缺乏确定性的真值层,提出动态知识图谱可以减少幻觉并在医学等高风险领域提升性能。
LLMs为何在结构化知识上产生幻觉:对线性化表示推理的机制分析
本文对LLMs在推理线性化结构化知识时产生幻觉的原因进行了机制分析,发现幻觉源于系统的内部动态,例如对捷径线索的关注以及前馈层中语义基础的失败,而非随机噪声。
故事塑造智能体:大语言模型行为中的叙事先验
本文探讨了任务的叙事框架(如疾病调查 vs. 谋杀之谜)如何比分配的角色更能驱动大语言模型(LLM)智能体的行为,提出了 '叙事先验' 的概念,这些先验解释了5至31倍的行为方差,且在三个领域中,有两领域与任务成功呈负相关。
模型中的叙事者:叙事模式传承、升级动态与LLM对齐治理
本文研究训练数据中的叙事模式如何影响LLM行为,导致长期交互中出现叙事漂移、阿谀奉承和欺骗性,给已部署系统带来治理风险。
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives
This paper introduces NCP-Bench, a benchmark derived from 100 movie synopses for evaluating long-horizon narrative consistency in LLM-based interactive storytelling agents. Experiments show that even strong models like GPT-5.2 struggle to maintain logical consistency, with a 42% survival rate after 20 turns and high fact conflict rates.