How LLMs Build Fictional Worlds: Setting and Narrative Space in AI-Generated Creative Storytelling

arXiv cs.CL Papers

Summary

This paper analyzes how Large Language Models (LLMs) use worldbuilding strategies in AI-generated creative storytelling, comparing them to human-authored fiction and finding that LLMs overproduce 'perceived space' while humans favor 'action space'.

arXiv:2609.02482v1 Announce Type: new Abstract: In this paper, we analyze how Large Language Models (LLMs) employ worldbuilding strategies, focusing on setting as one measurable dimension of storyworld construction. We compare 1,000 AI-generated stories per model in English and German with human-authored fiction from Project Gutenberg. Building on prior work, we operationalize setting through five types of narrative space: "action", "perceived," "visual," "descriptive" and "no space", identified using fine-tuned BERT classifiers for German and English. We generate narratives using GPT 4.1, LlaMA 3.3, Mistral 3.2, and Gemma 3 and compare their spatial distributions to a human-authored baseline. We find that human-authored texts predominantly employ "action space," grounding narratives in embodied character-environment interaction, whereas LLMs systematically overproduce "perceived space," emphasizing atmosphere and affect. This divergence remains stable across narrative time. Overall, our findings show that LLMs exhibit worldbuilding patterns that differ consistently from human-authored fiction in ways that are both model-specific and language-sensitive.
Original Article
View Cached Full Text

Cached at: 09/03/26, 05:54 AM

# How LLMs Build Fictional Worlds: Setting and Narrative Space in AI-Generated Creative Storytelling
Source: [https://arxiv.org/html/2609.02482](https://arxiv.org/html/2609.02482)
Katrin RohrbacherAffiliation:The Text and Language Lab, Department of Digital Humanities and Social Studies \(DHSS\), FAU Erlangen\-Nürnberg, GermanyEmail:[katrin\.rohrbacher@fau\.de](mailto:[email protected])Björn NiethAffiliation:Department Artificial Intelligence in Biomedical Engineering \(AIBE\), FAU Erlangen\-Nürnberg, GermanyAffiliation:Munich Center for Machine Learning \(MCML\), Munich, GermanyEmail:[bjoern\.nieth@fau\.de](mailto:[email protected])Emmanuelle SalinAffiliation:Department Artificial Intelligence in Biomedical Engineering \(AIBE\), FAU Erlangen\-Nürnberg, GermanyEmail:[emmanuelle\.salin@fau\.de](mailto:[email protected])Bjoern EskofierAffiliation:Department Artificial Intelligence in Biomedical Engineering \(AIBE\), FAU Erlangen\-Nürnberg, GermanyAffiliation:Chair of AI\-supported Therapy Decisions, LMU München, GermanyAffiliation:Munich Center for Machine Learning \(MCML\), Munich, GermanyAffiliation:Institute of AI for Health, Helmholtz Zentrum München, Neuherberg, GermanyEmail:[bjoern\.eskofier@fau\.de](mailto:[email protected])Michaela MahlbergAffiliation:The Text and Language Lab, Department of Digital Humanities and Social Studies \(DHSS\), FAU Erlangen\-Nürnberg, GermanyAffiliation:Department of Linguistics and Communication, University of Birmingham, United KingdomEmail:[michaela\.mahlberg@fau\.de](mailto:[email protected])

###### Abstract

In this paper, we analyze how Large Language Models \(LLMs\) employ worldbuilding strategies, focusing on setting as one measurable dimension of storyworld construction\. We compare 1,000 AI\-generated stories per model in English and German with human\-authored fiction from Project Gutenberg\. Building on prior work, we operationalize setting through five types of narrative space: “action,” “perceived,” “visual,” “descriptive” and “no space”, identified using fine\-tuned BERT classifiers for German and English\. We generate narratives using GPT 4\.1, LlaMA 3\.3, Mistral 3\.2, and Gemma 3 and compare their spatial distributions to a human\-authored baseline\. We find that human\-authored texts predominantly employ “action space,” grounding narratives in embodied character–environment interaction, whereas LLMs systematically overproduce “perceived space,” emphasizing atmosphere and affect\. This divergence remains stable across narrative time\. Overall, our findings show that LLMs exhibit worldbuilding patterns that differ consistently from human\-authored fiction in ways that are both model\-specific and language\-sensitive\.

## 1Introduction

Research in narratology, literary studies, and linguistics has identified worldbuilding as a central element of storytelling and one of its core communicative functions\. Herman argues that storyworlds are built from verbal and visual cues that prompt readers to mentally construct the depicted world as it evolves through interactions among characters, objects, and settings[Herman \(2002\)](https://arxiv.org/html/2609.02482#bib.bib20)\. In cognitive linguistics and corpus stylistics, researchers have explored how readers construct mental representations of fictional worlds drawing on “textual building blocks” or “world\-building elements”[Mahlberg \(2013\)](https://arxiv.org/html/2609.02482#bib.bib34);[Gavins \(2007\)](https://arxiv.org/html/2609.02482#bib.bib17)\. Gavins, in her work on “world\-building” in fiction, explicitly focuses on the opening paragraphs of stories, which serve as “an initial introduction to the fictional worlds about to be realized by the text”\([Gavins, 2007](https://arxiv.org/html/2609.02482#bib.bib17), p\. 133\)\. These approaches emphasize that fictional worlds emerge through specific textual structures, which collectively shape how readers experience and navigate storyworlds\.

Human\-authored

Lilian Boyd entered the small, rather shabby room, neat, though everything was well worn\.action spaceHer mother sat by a little work table busy with some muslin sewing and she looked up with a weary smile\.descriptive spaceLilian laid a five\-dollar bill on the table\.action space“Madame Lupton sails on Saturday,” she said\.no space“Oh how splendid it must be to go to Paris\!”no space

GPT 4\.1

Lilian Boyd entered the small, rather shabby room, neat, though everything was well worn\.action spaceA stuffy scent, reminiscent of dust and distant rain, lingered in the corners\.perceived space\[…\]Lilian closed the door behind her and allowed her eyes to sweep across the sparse furnishings\.visual spaceThe same old brown armchair squatted in the corner by the window\.descriptive space

Legend:actionperceiveddescriptivevisualno space

Figure 1:The same opening sentence \(bold\) continued by a human author \(Amanda M\. Douglas,*The Girls at Mount Morris*, 1914\) and GPT 4\.1, annotated by the setting classifier\.The task of creative text generation has received increasing attention in current LLM research\. Existing studies on LLMs’ capacities to generate fiction have primarily analyzed outputs through creativity tests or human judgments of stylistic quality[Zhao et al\. \(2024\)](https://arxiv.org/html/2609.02482#bib.bib51);[Ismayilzada et al\. \(2025\)](https://arxiv.org/html/2609.02482#bib.bib27)\. Other studies have examined specific narrative and stylistic features, such as embodied language, emotional expression, temporal structure, and plot construction, to better understand how LLMs generate narratives[Hicke et al\. \(2025\)](https://arxiv.org/html/2609.02482#bib.bib22);[Ishikawa and Yoshino \(2025\)](https://arxiv.org/html/2609.02482#bib.bib26);[Zhong et al\. \(2024\)](https://arxiv.org/html/2609.02482#bib.bib52);[Fatemi et al\. \(2024\)](https://arxiv.org/html/2609.02482#bib.bib15);[Ahuja et al\. \(2025\)](https://arxiv.org/html/2609.02482#bib.bib1)\. However, the spatial dimension of worldbuilding in AI\-generated fiction remains largely unexplored\.

In this study, we focus on setting as a central dimension of worldbuilding in creative storytelling, understanding it as one of the primary textual mechanisms \(alongside time and events\) through which fictional worlds are constructed\. We build on the framework introduced in[Rohrbacher \(2025b\)](https://arxiv.org/html/2609.02482#bib.bib39), which grounds narrative setting in the phenomenological notion oflived space[Hoffmann \(1978\)](https://arxiv.org/html/2609.02482#bib.bib24);[Ströker \(1965\)](https://arxiv.org/html/2609.02482#bib.bib43), i\.e\., space as it is experienced and inhabited by a perceiving subject, not as a neutral, abstract extension\. Lived space, in this sense, is not simply described but organized around characters’ perception and embodied engagement with their surroundings, conveying “what it is like” to inhabit a fictional world[Herman \(2009\)](https://arxiv.org/html/2609.02482#bib.bib21)\. The framework defines five categories\. Action space, perceived space, and visual space reflect distinct modes of spatial experience\. Descriptive space, by contrast, situates characters and objects in space without being anchored to any character’s point of view or experiential engagement\. “No space” covers sentences without spatial reference\.

To assess how LLMs use these spatial categories to construct narrative worlds, we apply the setting classifier introduced in[Rohrbacher \(2025b\)](https://arxiv.org/html/2609.02482#bib.bib39), a fine\-tuned BERT model originally developed for German fiction and extended in this study to English\. We generate 1,000 stories per model and language and compare the resulting distributions against a human\-authored corpus serving as a baseline\.

The contributions of this paper are as follows:

- •We find that LLM storytelling is systematically more atmospheric and less embodied/action\-driven than human fiction\.
- •Through a cross\-linguistic comparison of English and German corpora, we show that these deviation patterns are jointly shaped by model and language context\.
- •We release a large\-scale dataset of 8,000 AI\-generated stories \(1,000 per model across four LLMs and two languages\), matched with a human\-authored corpus\.
- •We release an English setting classifier and accompanying annotation resources that replicate an existing German classifier, supporting corpus\-scale, narratology\-informed evaluation of AI\-generated fiction\.

## 2Related Work

Despite the extensive narratological literature on how stories are told, relatively little machine learning research has drawn on narratological concepts to analyze how LLMs generate narratives\. Existing approaches to evaluating LLM\-generated narratives include creativity tests[Chakrabarty et al\. \(2024\)](https://arxiv.org/html/2609.02482#bib.bib8);[Ismayilzada et al\. \(2025\)](https://arxiv.org/html/2609.02482#bib.bib27), often adapted from psychological frameworks such as the Torrance Test of Creative Thinking[Zhao et al\. \(2024\)](https://arxiv.org/html/2609.02482#bib.bib51);[Marco et al\. \(2024\)](https://arxiv.org/html/2609.02482#bib.bib35);[Chen and Ding \(2023\)](https://arxiv.org/html/2609.02482#bib.bib10), which primarily measure human creative cognition rather than how creativity is realized in textual form\. Other work has used human raters, for instance to assess stylistic properties in order to infer perceived quality or humanlikeness[Chakrabarty et al\. \(2025\)](https://arxiv.org/html/2609.02482#bib.bib7), or to identify a story’s origin and rate its quality\([Sears and Weisberg, 2026](https://arxiv.org/html/2609.02482#bib.bib41)\)\. A third line of work focuses on identifying specific storytelling patterns in generated prose, including lexicon\- or dictionary\-based approaches for capturing concepts such as embodiment in narrative[Hicke et al\. \(2025\)](https://arxiv.org/html/2609.02482#bib.bib22)\. However, such approaches struggle with polysemy and contextual meaning, a particular challenge in fiction, where the same word may carry spatial or non\-spatial meaning depending on context\. Our classifier, trained on manually annotated prose, implicitly learns contextual meaning rather than relying on surface vocabulary\.

To model patterns of setting, we draw on prior work demonstrating that fine\-tuning transformer models on hand\-annotated data is effective for capturing complex narratological phenomena, including narrativity, events, and space[Antoniak et al\. \(2024\)](https://arxiv.org/html/2609.02482#bib.bib2);[Vauth et al\. \(2021\)](https://arxiv.org/html/2609.02482#bib.bib48);[Kababgi et al\. \(2024\)](https://arxiv.org/html/2609.02482#bib.bib28);[Soni et al\. \(2023\)](https://arxiv.org/html/2609.02482#bib.bib42)\. Unlike dictionary\-based methods, transformer classifiers capture contextual and long\-range semantic dependencies, allowing for the detection of abstract patterns beyond n\-gram overlap[Vaswani et al\. \(2023\)](https://arxiv.org/html/2609.02482#bib.bib47);[Wankmüller \(2024\)](https://arxiv.org/html/2609.02482#bib.bib50)\.

Research on long\-form narrative generation remains limited\. Existing studies often focus on short passages or synopses[Tian et al\. \(2024\)](https://arxiv.org/html/2609.02482#bib.bib45)\.111Here, we uselong\-formto refer to multi\-chapter narratives of several thousand words, as opposed to single passages, synopses, or summaries\.While LLMs continue to struggle with extended narrative coherence, longer context windows and improved prompting strategies have made extended generation more feasible for recent models \(e\.g\.,[Wang et al\. 2025](https://arxiv.org/html/2609.02482#bib.bib49);[Bae and Kim 2024](https://arxiv.org/html/2609.02482#bib.bib3)\)\.

## 3Method

### 3\.1Dataset

We draw a random sample of 1,000 works per language from English and German fiction corpora of approximately 4,300 public\-domain texts\. The English corpus \(1800–1920\) was scraped from Project Gutenberg based on the metadata provided by[Cuthbert et al\. \(2019\)](https://arxiv.org/html/2609.02482#bib.bib13)\. Genre labels for the English corpus were assigned by us using the subject and topic metadata that Project Gutenberg supplies for each work\. The German corpus \(1780–1940\) was drawn from[Rohrbacher \(2025a\)](https://arxiv.org/html/2609.02482#bib.bib38), itself sourced primarily from Projekt Gutenberg\-DE with a small fraction coming from German works in the English Project Gutenberg\. Genre labels for the German corpus were taken from Projekt Gutenberg\-DE, which follows the categories of the German book trade, and are provided with the corpus\. Both corpora consist primarily of novels and novellas and cover a broad range of narrative subgenres, including speculative fiction, crime fiction, fairy tales, and young adult literature\.

### 3\.2Narrative generation

We use prompts that contain the first sentence of each human\-written text to generate long\-form narratives of four chapters averaging 6,500–9,400 words per model \(Appendix[I](https://arxiv.org/html/2609.02482#A9)\)\. We refer to these model\-generated texts as “continuations”\. Genre information is provided as an additional conditioning input \(e\.g\., novel, speculative fiction, fairy tale\)\. To analyze variation across model families, we prompted three open\-source models and one closed\-source model \(LlaMA\-3\.3 70B[Grattafiori et al\. \(2024\)](https://arxiv.org/html/2609.02482#bib.bib19), Mistral 3\.2 24B222[https://huggingface\.co/mistralai/Mistral\-Small\-3\.2\-24B\-Instruct\-2506](https://huggingface.co/mistralai/Mistral-Small-3.2-24B-Instruct-2506), Gemma 3 27B[Gemma Team \(2025\)](https://arxiv.org/html/2609.02482#bib.bib18), and GPT 4\.1 \(gpt\-4\.1\-2025\-04\-14\)333[https://openai\.com/index/gpt\-4\-1/](https://openai.com/index/gpt-4-1/)\)\.444For brevity, we henceforth refer to these models as LlaMA 3\.3, Mistral 3\.2, Gemma 3, and GPT 4\.1 respectively\.Model selection was informed by prior evaluations of creative writing performance[Paech \(2023\)](https://arxiv.org/html/2609.02482#bib.bib37)\. Additional models were piloted but excluded due to recurrent generation failures, including truncation, repetition, and unintended language switching\.555Excluded models: Mixtral, LlaMA\-3\-8B, QwQ\-32B, Qwen\-2\.5, Apertus\.We sampled all models with a temperature and top\-p of 1 to study model behavior without the influence of sampling parameters, making results comparable across models\. For reproducibility, we used a seeded random function with a fixed seed\.

Table 1:Average sentence length in words for English \(en\) and German \(de\) texts\. The matched human excerpts average 10,761 words in English and 7,953 in German\. Story lengths for the generated texts are reported in Appendix[I](https://arxiv.org/html/2609.02482#A9)Because long\-form generation remains challenging for many models, particularly open\-source ones, we employ an iterative, chapter\-based prompting strategy\. Rather than issuing a single prompt, we guide generation incrementally, with each chapter conditioned on the model’s prior output\. As shown in Table[2](https://arxiv.org/html/2609.02482#S3.T2), the model is assigned a system role and given stepwise user instructions for generating the narrative chapters\. Prompts were refined through iterative testing to suppress meta\-commentary, user\-directed dialogue, repetition, and premature endings\. The resulting dataset, a corpus of 8,000 AI\-generated stories \(1,000 per model for each language\), is available in our code repository, together with the metadata and opening sentences of the matched human\-authored sample\.666Code and data used in this project can be found here:[https://github\.com/BjoernNieth/worldbuilding\-AI](https://github.com/BjoernNieth/worldbuilding-AI)\. Dataset details are reported in Table[1](https://arxiv.org/html/2609.02482#S3.T1)\.

We report four robustness checks\. To rule out memorization, we measured 13\-gram overlap between generated texts and their source books\. To confirm robustness to prompt formulation, we ran three prompt variants for each open\-source model and language, which vary phrasing and verbosity \(see Appendix[C](https://arxiv.org/html/2609.02482#A3)\)\. Because each chapter is prompted separately, every chapter start may behave like a story opening and drive the patterns we report across narrative time\. To test this, we re\-bin the generated texts by position within chapter \(Section[3\.4](https://arxiv.org/html/2609.02482#S3.SS4)\)\. Finally, we include genre as a covariate in the statistical model, since it is supplied to the models during generation and could influence the spatial distributions \(Section[3\.5](https://arxiv.org/html/2609.02482#S3.SS5)\)\.

Table 2:Instructions for model\-generated narratives\.\{genre\}and\{sentence\}are filled with metadata and the opening sentence of each source text\. Generation was capped at 3,000 tokens per chapter and stopped at the end\-of\-sequence token, for a total of 10,000–12,000 tokens per story \(Appendix[I](https://arxiv.org/html/2609.02482#A9)\)\.
### 3\.3Setting classification

We use a fine\-tuned BERT classifier that assigns each sentence to one of five categories: action space, perceived space, visual space, descriptive space, and “no space”[Rohrbacher \(2025b\)](https://arxiv.org/html/2609.02482#bib.bib39)\. The five categories capture distinct modes of spatial representation\. See Table[3](https://arxiv.org/html/2609.02482#S3.T3)for full definitions and examples\. The classifier was originally developed for German\-language fiction\. For the present study, we replicate it for English by fine\-tuning on a manually annotated dataset of English fictional prose following the same annotation scheme, achieving comparable performance \(see Appendix[A\.2](https://arxiv.org/html/2609.02482#A1.SS2)\)\. Since the classifier was fine\-tuned on human\-authored fictional prose, we validated its generalization to AI\-generated text using a manually annotated sample \(see Section[4\.1](https://arxiv.org/html/2609.02482#S4.SS1)\)\.

Table 3:Categories of narrative setting used by the setting classifier \(adapted from[Rohrbacher \(2025a\)](https://arxiv.org/html/2609.02482#bib.bib38)\)\. Labels follow English translations of the original German terms:Aktionsraum\(action space\),gestimmter Raum\(perceived space\),Anschauungsraum\(visual space\)\.
### 3\.4Analysis design

To compare how narrative environments are structured across AI\-generated and human\-authored texts, we compute the proportion of each category across all sentences in a corpus\. We apply the classifier to both, matching each literary excerpt to the length of its corresponding generated story\.777[Lucy and Bamman \(2021\)](https://arxiv.org/html/2609.02482#bib.bib33)adopt a comparable design, prompting GPT\-3 with opening sentences from novels and comparing the output against length\-matched excerpts from the same books\.We analyze setting at two levels of granularity\. At the micro level, we focus on the opening of the text\. Narratological research identifies openings as the primary site of world\-establishment[Gavins \(2007\)](https://arxiv.org/html/2609.02482#bib.bib17);[Herman \(2009\)](https://arxiv.org/html/2609.02482#bib.bib21), but they have no fixed textual boundary\. We therefore operationalize the opening as the first 15 sentences of each text, an approximation we adopt for the present analysis\. At the macro level, we segment texts into ten equal\-length sections and track the normalized frequency of each category across these intervals\. Here, normalized frequency refers to the proportion of sentences assigned to a given category within each section\. This design is motivated by prior work showing that fictional texts follow recognizable patterns in how setting unfolds across narrative time[Rohrbacher \(2025b\)](https://arxiv.org/html/2609.02482#bib.bib39);[Boyd et al\. \(2020\)](https://arxiv.org/html/2609.02482#bib.bib4)\. Although the generated excerpts lack canonical endings, this normalization allows comparison across texts of varying length\.

Because narratives were generated chapter by chapter \(Section[3\.2](https://arxiv.org/html/2609.02482#S3.SS2)\), openings recur at every chapter boundary, and the ten\-section grid is not aligned to them\. We therefore additionally re\-bin each generated text by position within chapter, dividing each of the four chapters into quarters \(16 bins per story\)\.

### 3\.5Statistical analysis

To test whether setting category distributions differ systematically between human\-authored and model\-generated texts across narrative time, we fit a Generalized Linear Mixed Model \(GLMM\) for each spatial category using the glmmTMB R package[Brooks et al\. \(2017\)](https://arxiv.org/html/2609.02482#bib.bib6)\. We modeled each setting category as a function of a five\-level author factor \(Human, GPT 4\.1, Gemma 3, LlaMA 3\.3, Mistral 3\.2\) crossed with narrative section\. To account for story\-specific variation, the model included a by\-story random intercept\. Genre was included as a covariate\. Due to the small size of most genre types, we code genre as novel vs\. other fiction to avoid modeling minor subgenres separately\. Each generated story takes the genre label of the source text that seeded it, so the genre distribution is identical across all five author conditions \(English: 69\.2% novel; German: 79\.6%\)\. The full specification is space ~ author×\\timessection \+ genre \+ \(1 \| story\)\. As each outcome is a proportion bounded in\[0,1\]\[0,1\], including exact zeros and ones, we used the ordbeta family[Kubinec \(2023\)](https://arxiv.org/html/2609.02482#bib.bib29)\. Significance was assessed via Type III Waldχ2\\chi^\{2\}tests\. Post\-hoc contrasts were computed as estimated marginal means usingemmeans[Lenth and Piaskowski \(2026\)](https://arxiv.org/html/2609.02482#bib.bib32)\.

## 4Results

### 4\.1Classifier validation

Two annotators \(the first author and a research student\) manually labeled a stratified sample of 600 sentences drawn from the AI\-generated texts, covering all four models and both languages\. Inter\-annotator agreement was substantial \(κ=0\.736\\kappa=0\.736, 79\.2% raw agreement\)\. The classifier achieved an accuracy of 78\.0% and a mean Cohen’sκ\\kappaof 0\.725, closely approaching the human inter\-annotator ceiling\. Performance was consistent across languages and models, with the exception of Mistral 3\.2 \(see Appendix[A\.3](https://arxiv.org/html/2609.02482#A1.SS3)\)\.

### 4\.2Analysis of setting in story openings

Figure[2](https://arxiv.org/html/2609.02482#S4.F2)shows the normalized frequency of each spatial category in the opening 15 sentences, compared across models and languages\. It reveals a pronounced contrast between human\-authored and model\-generated openings\. Perceived space is the most frequent category in both model\-generated texts and human\-authored fiction\. Whereas human\-authored fiction shows normalized frequencies of around 0\.19 in both languages, all four models exceed this level, with GPT 4\.1 reaching around 0\.47 in English and 0\.38 in German, roughly twice the human baseline\. For action space, Gemma 3 differs most strongly from the human baseline, while the other three models remain broadly within the same range\. Visual space shows a moderate but consistent increase in model\-generated text, most pronounced for LlaMA 3\.3, which reaches roughly twice the human baseline in both languages\. Descriptive space is the category that models reproduce most closely, with language\-specific differences shown in both human and AI\-generated texts\. The overall gap between human\-authored and generated texts is more pronounced in English than in German\.

Human\-authored Frank Norris, McTeague: A Story of San Francisco \(1899\)

It was Sunday, and, according to his custom on that day, McTeague took his dinner at two in the afternoon at the car conductors’ coffee\-joint on Polk Street\.action spaceHe had a thick gray soup; heavy, underdone meat, very hot, on a cold plate; two kinds of vegetables; and a sort of suet pudding, full of strong butter and sugar\.perceived spaceOn his way back to his office, one block above, he stopped at Joe Frenna’s saloon and bought a pitcher of steam beer\.action spaceIt was his habit to leave the pitcher there on his way to dinner\.action spaceOnce in his office, he took off his coat and shoes, unbuttoned his vest, and, having crammed his little stove full of coke, lay back in his operating chair at the bay window, reading the paper, drinking his beer, and smoking his huge porcelain pipe while his food digested; crop\-full, stupid, and warm\.action spaceBy and by, gorged with steam beer, and overcome by the heat of the room, the cheap tobacco, and the effects of his heavy meal, he dropped off to sleep\. \[…\]perceived space

GPT 4\.1 generated Continuation from same opening sentence

It was Sunday, and, according to his custom on that day, McTeague took his dinner at two in the afternoon at the car conductors’ coffee\-joint on Polk Street\.action spaceThe sky outside was a sullen, unpolished grey, holding back the barely audible grumble of distant thunder; a few straggling shoppers hurried past the big front window, bent under coats and unhappy hats, all the world’s machinery seeming to slow with the prospect of Sunday’s late rain\.perceived spaceInside, under an uneven circle of yellow light, McTeague sat waiting, his big hands turning and returning a battered steel fork, above a square wooden patch of table darkened by years of elbows and spilled drinks\.perceived spaceThe place was narrow, long, heavy with the paraffin tang of old oilcloth and the low murmurs of the regulars, conductors in soiled blue coats, a few shop clerks, a woman with a battered net of oranges, all as familiar and faded to McTeague as the blurred wallpaper\.perceived spaceThere were no clocks in the coffee\-joint\. \[…\]descriptive space

Legend:action spaceperceived spacedescriptive space

Table 4:Comparison of story openings: human\-authored passage from Frank Norris’*McTeague: A Story of San Francisco*\([Norris, 1899](https://arxiv.org/html/2609.02482#bib.bib36)\)versus GPT 4\.1 continuation of the same opening sentence\. Colours indicate the narrative space type assigned by the setting classifier\.Figure 2:Normalized frequencies of setting categories in the opening sentences \(first 15 sentences\) of human\-authored and AI\-generated texts in English and German\. Error bars indicate 95% confidence intervals\.The example in Table[4](https://arxiv.org/html/2609.02482#S4.T4)illustrates how these distributional differences operate at the sentence level\. After the shared opening sentence, the two passages immediately diverge\. Norris constructs space primarily through movement and touch\. The character McTeague traverses a habitual route from restaurant to saloon to office, and the setting emerges through that trajectory\. Objects serve functional roles, the pitcher left at the saloon to be collected on the way back, the stove crammed with coke\. Perceived space, where it appears, is bound to the body and conveys temperature and smoke only through their effect on the character\. GPT 4\.1, by contrast, arrests movement from the second sentence onward\. Three consecutive sentences of perceived space follow\. Objects become prominent as sensory props rather than instruments for purposeful action, the “turning and returning” of the fork, the table textured by “years” of use\. The environment is rendered quasi\-anthropomorphic, its grey sky “sullen” and “the world’s machinery seeming to slow\.” Unlike in Norris’ passage, the detail of the external world is emphasized\. The passage illustrates a storyworld pervaded by atmosphere, where the environment presses in on the character\. What the classifier labels as perceived space is, at the level of narrative experience, a world felt before it is actively inhabited\.

### 4\.3Temporal distribution of setting across narrative time

#### 4\.3\.1Overall differences in proportion

In human\-authored fiction, action space is the most frequently produced spatial category\. LLMs deviate systematically from the distributional profile of human\-authored fiction, most consistently in their overuse of perceived space\. Figure[3](https://arxiv.org/html/2609.02482#S4.F3)shows the deviation of each model from the human baseline across ten narrative sections for English and German\. The statistical model confirms significant model differences in the first section and a significant model×\\timessection interaction, indicating that models diverge from human authors in ways that vary across narrative time \(English: model main effectχ2​\(4\)=2509\.37\\chi^\{2\}\(4\)=2509\.37,p<\.001p<\.001; model×\\timessection interactionχ2​\(36\)=2013\.48\\chi^\{2\}\(36\)=2013\.48,p<\.001p<\.001\)\. Post\-hoc contrasts confirm that all four LLMs produce significantly more perceived space than human authors across every narrative section \(allp<\.001p<\.001after Holm correction\)\. GPT 4\.1 shows the largest deviations on average \(∼0\.14\{\\sim\}0\.14–0\.250\.25above baseline\) and LlaMA 3\.3 the second largest \(∼0\.10\{\\sim\}0\.10–0\.230\.23\), while Gemma 3 remains closest to the human baseline \(∼0\.06\{\\sim\}0\.06–0\.100\.10\)\. For action space, GPT 4\.1 is the only model that remains close to or above the human baseline\. Gemma 3 and Mistral 3\.2 fall below it from the opening section onward, while LlaMA 3\.3 diverges from section 2 onward\. These differences produce a significant model effect \(English:χ2​\(4\)=456\.78\\chi^\{2\}\(4\)=456\.78,p<\.001p<\.001\)\. LlaMA 3\.3 shows the largest deficit overall, based on post\-hoc contrasts \(estimates−0\.04\-0\.04–−0\.07\-0\.07across sections 2–10, allp<\.001p<\.001\)\. This pattern is reflected in the all\-space aggregate: GPT 4\.1 produces a consistently higher proportion of spatially marked sentences than human authors\.

Figure 3:Deviation of model predictions from the human baseline for perceived and action space across narrative sections, comparing English and German corpora\.
#### 4\.3\.2Differences in trend

Figure[4](https://arxiv.org/html/2609.02482#S4.F4)shows action and descriptive space frequency across ten narrative sections\. For action space, GPT 4\.1 tracks closely with the human baseline and shares its slight upward trend\. LlaMA 3\.3 declines continuously, while Gemma 3 and Mistral 3\.2 stabilize below the baseline after an initial drop\. For descriptive space, all models broadly follow the human baseline’s declining trend, with GPT 4\.1 and Gemma 3 remaining somewhat elevated, and Mistral 3\.2 and LlaMA 3\.3 falling below it\. For perceived space \(Figure[9](https://arxiv.org/html/2609.02482#A4.F9)in Appendix[D](https://arxiv.org/html/2609.02482#A4)\), the human baseline starts moderately and remains largely flat across narrative sections, whereas all models start higher and show no comparable stabilization\. GPT 4\.1 and LlaMA 3\.3 in particular fluctuate considerably throughout\. Part of this variance is attributable to chapter boundaries in the generation procedure \(Section[4\.4](https://arxiv.org/html/2609.02482#S4.SS4), Appendix[H](https://arxiv.org/html/2609.02482#A8)\)\. Beyond the boundary effect, the models distribute atmospheric density less evenly across a narrative than human authors do\.

Figure 4:Normalized frequency ofaction spaceanddescriptive spaceacross narrative sections for human\-authored and AI\-generated texts in English\. Shaded bands indicate±\\pm1 standard error\.
#### 4\.3\.3Comparison English vs\. German

The cross\-linguistic comparison \(Figure[3](https://arxiv.org/html/2609.02482#S4.F3)\) shows a reversal for perceived space\. While all models exceed the English human baseline, most fall at or below the German baseline \(GPT 4\.1 excepted, which remains elevated in both languages\)\. Post\-hoc contrasts show that Gemma 3 and Mistral 3\.2 are indistinguishable from the German human baseline in several later sections, while LlaMA 3\.3 falls consistently below it from section 3 onward \(χ2​\(36\)=1860\.15\\chi^\{2\}\(36\)=1860\.15,p<\.001p<\.001\)\. For action space, the pattern of deviation replicates across languages but is more pronounced in German \(χ2​\(36\)=1449\.79\\chi^\{2\}\(36\)=1449\.79,p<\.001p<\.001\)\. LlaMA’s declining trajectory and Gemma’s and Mistral’s stable deficit are visible in both corpora, with the gap widening in German\. The all\-space aggregate mirrors this\. GPT 4\.1 exceeds the human baseline in both languages, whereas the remaining models, which approach human levels in English, fall clearly below it in German\.

### 4\.4Ablation studies

##### Memorization\.

Across both languages we found no instances of verbatim reconstruction, with the single exception of Gemma 3 reproducing the well\-known opening sentence of Dickens’A Tale of Two Cities\.

##### Prompt sensitivity\.

As shown in Figures[7](https://arxiv.org/html/2609.02482#A3.F7)and[8](https://arxiv.org/html/2609.02482#A3.F8), prompt formulation influences the absolute frequency of spatial categories to some degree, but overall trajectories across narrative sections remain consistent across all conditions, confirming that the main findings are not an artifact of the specific prompt used\.

##### Chapter\-boundary position\.

Perceived space is elevated at the start of every chapter relative to the second quarter of the chapter, across all eight model\-language conditions, and rises again in the final quarter in English \(except Gemma 3\) and, in German, for GPT 4\.1\. The effect is small relative to the difference we report\. The largest within\-chapter range is 0\.09, whereas the English models exceed the human baseline by 0\.11 to 0\.21 overall, and even at their within\-chapter minimum all four produce two to three times as much perceived space as human authors\. The overproduction is moreover already present in chapter 1, which is generated from a single prompt before any continuation instruction\. Part of the temporal variance described in Appendix[D](https://arxiv.org/html/2609.02482#A4)is therefore attributable to the chapter\-based prompting strategy, but the overproduction itself is not\.

##### Genre\.

Genre is significant for perceived and action space in English \(χ2​\(1\)=18\.63\\chi^\{2\}\(1\)=18\.63and23\.1423\.14, bothp<\.001p<\.001\) and for all categories except descriptive space in German\. The author main effect and the author×\\timessection interaction remain highly significant with genre included \(Appendix[G](https://arxiv.org/html/2609.02482#A7)\)\.

## 5Discussion

In human\-authored fiction, action space in story openings anchors readers in sequences of habitual, embodied action, generating narrative momentum through characters’ movement and routine interaction with space\. LLM\-generated openings, instead, privilege perceived space, foregroundingStimmung\(i\.e\., mood or atmosphere\) over action, resulting in a diffuse “background feeling”[Colombetti \(2013\)](https://arxiv.org/html/2609.02482#bib.bib12)untethered from what characters actually do\.888The original German term isgestimmter Raum, i\.e\. space as it reflects atmosphere or mood, which more directly captures the connotation of space as permeated by feeling, a nuance that the English “perceived space” only partially conveys\.This skew contributes to an empirically recognizable generative AI style—rich in affect and ambience, but less grounded in embodied action and spatial concreteness\. This pattern aligns with prior work documenting LLMs’ preference for sensorial and affect\-laden language[Hicke et al\. \(2025\)](https://arxiv.org/html/2609.02482#bib.bib22);[Lee and Lim \(2024\)](https://arxiv.org/html/2609.02482#bib.bib31), and extends it by showing how these preferences manifest at the level of narrative setting and shape storyworld construction\.

A plausible explanation is that post\-training alignment methods might amplify this tendency\. If human raters reward emotionally resonant, atmospheric prose, generation may be biased systematically toward perceived space\. Importantly, the pattern we observe does not appear to be driven by historical periodicity in the human baseline\. Spatial category distributions in the human\-authored corpora show no systematic directional trend across the main period of the corpus \(see Appendix[F](https://arxiv.org/html/2609.02482#A6)\)\. This suggests the observed differences reflect a systematic contrast between human and LLM writing rather than a temporal drift in narrative space composition\.

A second explanation concerns the generation procedure itself\. Because each chapter is prompted afresh, every chapter start is treated as an opening, and perceived space rises at chapter boundaries relative to chapter middles \(Appendix[H](https://arxiv.org/html/2609.02482#A8)\)\. This accounts for part of the fluctuation across narrative sections, but not for the overall level, which is already elevated in chapter 1\. Chapter\-level prompting is a necessity given current long\-form performance, and a different strategy would produce a different boundary profile without removing the underlying skew\.

The cross\-linguistic asymmetry suggests that these tendencies are not uniformly stable across languages\. Open\-source models align more closely with GPT 4\.1 in English than in German\. In the latter, their perceived space distributions converge toward or fall below the human baseline, and their action space deficit widens\. The perceived and action spaces of GPT 4\.1, by contrast, remain elevated in both languages\. In most cases, the three open\-source models show comparatively little variation relative to each other despite differing in architecture and size\. This suggests that the observed spatial tendencies reflect shared properties of instruction\-tuned LLMs more broadly, such as preference optimization or training data composition, rather than model\-specific factors\.

Included as a covariate, genre has a significant effect for several categories in both languages \(Appendix[G](https://arxiv.org/html/2609.02482#A7)\)\. But because genre is matched across author conditions, this does not affect the human\-model comparison\. It does, however, indicate that spatial composition varies with genre, and the framework could be applied to genre\-stratified corpora to examine this directly\.

The fact that spatial category distributions alone are sufficient to identify the generating model well above chance \(see Appendix[E](https://arxiv.org/html/2609.02482#A5)\) indicates that setting constitutes a reliable stylistic marker of AI\-generated fiction\. We identify two directions that could build on this\. The first is space\-conditioned generation, in which models would be explicitly constrained to produce texts belonging to certain narrative spaces during prompting, which could mitigate the differences we report\. The second is empirical validation with readers\. Our analysis documents a difference in spatial composition, which is manifest in the texts themselves, but does not establish what that difference means for the reading experience\. Existing work compares human\-authored and AI\-generated stories as wholes and finds that readers rate the AI\-generated ones as more absorbing\([Sears and Weisberg, 2026](https://arxiv.org/html/2609.02482#bib.bib41)\), but these judgements are not linked to any measured property of the texts\. A reader study could vary spatial composition directly and test whether it corresponds to differences in immersion or perceived quality\.

## 6Conclusion

We introduced a narratology\-informed framework for analyzing worldbuilding in AI\-generated fiction, comprising five spatial categories grounded in narrative theory, a fine\-tuned BERT classifier for English \(extending an existing German classifier\), and a corpus\-scale application across four LLMs in both languages\. The results reveal consistent differences between LLM\-generated and human\-authored text in how narrative space is constructed\. LLMs systematically overrepresent perceived space while producing less action space than human authors, a pattern that holds across models and across narrative sections, though with model\-specific and cross\-linguistic variation in magnitude\. GPT 4\.1 shows the largest and most consistent overproduction of perceived space across both languages, while remaining close to the human baseline for action and descriptive space\. These findings suggest that LLMs do not reproduce the spatial distributions of human fiction, but construct storyworlds in a way that differs systematically from literary norms\. The spatial metrics that we propose in this paper, derived from narratological theory and grounded in literary scholarship, offer a path toward richer evaluation of AI\-generated narratives\.

## Limitations

Long\-form generation remains challenging for LLMs, and the patterns we report, particularly regarding narrative time, may therefore reflect tendencies rather than fully stable narrative strategies\. Because we prompted the models to produce “chapters”, part of the variance in spatial distributions across narrative time is attributable to chapter boundaries, where perceived space is elevated \(Section[4\.4](https://arxiv.org/html/2609.02482#S4.SS4), Appendix[H](https://arxiv.org/html/2609.02482#A8)\)\. The classifier’s comparatively lower performance on Mistral 3\.2 outputs, most pronounced in German, means that the results for this model should be interpreted with some caution\. Additionally, our analysis focuses on English and German texts from public\-domain Project Gutenberg \(English: 1800\-\-1920; German: 1780\-\-1940\), which may limit the generalizability of our findings to contemporary fiction and to languages beyond English and German\. Extending the analysis to contemporary fiction remains an open direction, and would show whether the patterns we report hold for present\-day fiction\. Since we use only the first sentence of each human\-authored text as a prompt and provide no information on year of publication, we cannot expect LLMs to reproduce the stylistic conventions of this historical period\.999Recent research has shown that LLMs struggle to reproduce the “period style”[Underwood et al\. \(2025\)](https://arxiv.org/html/2609.02482#bib.bib46)of fictional prose when specifically prompted with examples from that period\.Our coarse novel vs\. other coding cannot capture genre\-specific worldbuilding strategies \(e\.g\., “speculative fiction” vs\. “fairy tales”\), and subgenre\-stratified sample sizes are too small relative to the dominant novel and novella categories to test whether the divergence we report varies in magnitude across genres\. A genre\-balanced corpus would be needed to answer this, which we leave to future work\. Despite these limitations, our results provide valuable insight into systematic differences between AI\-generated and human\-authored fiction\.

## Acknowledgments

The authors gratefully acknowledge the scientific support and HPC resources provided by the Erlangen National High Performance Computing Center \(NHR@FAU\) of the Friedrich\-Alexander\-Universität Erlangen\-Nürnberg \(FAU\)\. The hardware is partially funded by the German Research Foundation \(DFG\)\. The authors further acknowledge support from the Alexander von Humboldt Foundation as part of the Alexander von Humboldt Professorship endowed by the German Federal Ministry of Research, Technology and Space \(BMFTR\)\. The authors thank Jan\-Oliver Reincke for help with the annotation\.

## References

- Ahuja et al\. \(2025\)Kabir Ahuja, Melanie Sclar, and Yulia Tsvetkov\. 2025\.Finding flawed fictions: Evaluating complex reasoning in language models via plot hole detection\.*Preprint, arXiv:2504\.11900*\.
- Antoniak et al\. \(2024\)Maria Antoniak, Joel Mire, Maarten Sap, Elliott Ash, and Andrew Piper\. 2024\.[Where do people tell stories online? story detection across online communities](https://doi.org/10.18653/v1/2024.acl-long.383)\.In*Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 7104–7130, Bangkok, Thailand\. Association for Computational Linguistics\.
- Bae and Kim \(2024\)Minwook Bae and Hyounghun Kim\. 2024\.[Collective critics for creative story generation](https://doi.org/0.18653/v1/2024.emnlp-main.1046)\.In*Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing*, pages 18784–18819, Miami, Florida, USA\. Association for Computational Linguistics\.
- Boyd et al\. \(2020\)Ryan L Boyd, Kate G Blackburn, and James W Pennebaker\. 2020\.The narrative arc: Revealing core narrative structures through text analysis\.*Science advances*, 6\(32\):eaba2196\.
- Brontë \(1847\)Emily Brontë\. 1847\.[*Wuthering Heights*](https://www.gutenberg.org/ebooks/768)\.Project Gutenberg\.
- Brooks et al\. \(2017\)Mollie E\. Brooks, Kasper Kristensen, Koen J\. van Benthem, Arni Magnusson, Casper W\. Berg, Anders Nielsen, Hans J\. Skaug, Martin Mæchler, and Benjamin M\. Bolker\. 2017\.[glmmTMB balances speed and flexibility among packages for zero\-inflated generalized linear mixed modeling](https://doi.org/10.32614/RJ-2017-066)\.*The R Journal*, 9\(2\):378–400\.
- Chakrabarty et al\. \(2025\)Tuhin Chakrabarty, Jane C\. Ginsburg, and Paramveer Dhillon\. 2025\.[Readers Prefer Outputs of AI Trained on Copyrighted Books over Expert Human Writers](https://doi.org/10.2139/ssrn.5606570)\.
- Chakrabarty et al\. \(2024\)Tuhin Chakrabarty, Philippe Laban, Divyansh Agarwal, Smaranda Muresan, and Chien\-Sheng Wu\. 2024\.[Art or Artifice? Large Language Models and the False Promise of Creativity](https://doi.org/10.48550/arXiv.2309.14556)\.*arXiv preprint*\.ArXiv:2309\.14556 \[cs\]\.
- Chapman \(1910\)Allen Chapman\. 1910\.[*Tom Fairfield’s Hunting Trip, or Lost in the Wilderness*](https://www.gutenberg.org/ebooks/44457)\.Project Gutenberg\.
- Chen and Ding \(2023\)Honghua Chen and Nai Ding\. 2023\.[Probing the “Creativity” of Large Language Models: Can models produce divergent semantic association?](https://doi.org/10.18653/v1/2023.findings-emnlp.858)In*Findings of the Association for Computational Linguistics: EMNLP 2023*, pages 12881–12888, Singapore\. Association for Computational Linguistics\.
- Cicchetti and Feinstein \(1990\)Domenic V\. Cicchetti and Alvan R\. Feinstein\. 1990\.[High agreement but low kappa: II\. resolving the paradoxes](https://doi.org/10.1016/0895-4356(90)90159-M)\.*Journal of Clinical Epidemiology*, 43\(6\):551–558\.
- Colombetti \(2013\)Giovanna Colombetti\. 2013\.[*The Feeling Body: Affective Science Meets the Enactive Mind*](http://site.ebrary.com/id/11204352)\.MIT Press, Cambridge, MA\.
- Cuthbert et al\. \(2019\)Michael Scott Asato Cuthbert, Lisa Tagliaferri, Stephan Risi, and 1 others\. 2019\.[Computational reading of gender in novels, 1770–1922](http://gendernovels.digitalhumanitiesmit.org/)\.MIT Digital Humanities Lab\.
- Devlin et al\. \(2019\)Jacob Devlin, Ming\-Wei Chang, Kenton Lee, and Kristina Toutanova\. 2019\.[BERT: Pre\-training of Deep Bidirectional Transformers for Language Understanding](https://doi.org/10.18653/v1/N19-1423)\.In*Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long and Short Papers\)*, pages 4171–4186, Minneapolis, Minnesota\. Association for Computational Linguistics\.
- Fatemi et al\. \(2024\)Bahare Fatemi, Mehran Kazemi, Anton Tsitsulin, Karishma Malkan, Jinyeong Yim, John Palowitch, Sungyong Seo, Jonathan Halcrow, and Bryan Perozzi\. 2024\.Test of time: A benchmark for evaluating llms on temporal reasoning\.*arXiv preprint arXiv:2406\.09170*\.
- Garland \(1915\)John Garland\. 1915\.[*Ross Grant Tenderfoot*](https://www.gutenberg.org/ebooks/34296)\.Project Gutenberg\.
- Gavins \(2007\)Joanna Gavins\. 2007\.*Text world theory : an introduction*\.Edinburgh University Press, Edinburgh\.OCLC: 1162422251\.
- Gemma Team \(2025\)Gemma Team\. 2025\.[Gemma 3 technical report](https://arxiv.org/abs/2503.19786)\.ArXiv:2503\.19786 \[cs\.CL\]\.
- Grattafiori et al\. \(2024\)Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al\-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, and 542 others\. 2024\.[The Llama 3 Herd of Models](https://doi.org/10.48550/arXiv.2407.21783)\.*arXiv preprint*\.ArXiv:2407\.21783 \[cs\]\.
- Herman \(2002\)David Herman\. 2002\.*Story Logic: Problems and Possibilities of Narrative*\.University of Nebraska Press, Lincoln and London\.
- Herman \(2009\)David Herman\. 2009\.*Basic elements of narrative*\.Wiley\-Blackwell, Chichester, U\.K\. ; Malden, MA\.OCLC: 229467488\.
- Hicke et al\. \(2025\)Rebecca M\. M\. Hicke, Sil Hamilton, and David Mimno\. 2025\.[The Zero Body Problem: Probing LLM Use of Sensory Language](https://doi.org/10.48550/arXiv.2504.06393)\.*arXiv preprint*\.ArXiv:2504\.06393 \[cs\]\.
- Hill \(1916\)Grace Brooks Hill\. 1916\.[*The Corner House Girls on a Houseboat*](https://www.gutenberg.org/ebooks/38609)\.Project Gutenberg\.
- Hoffmann \(1978\)Gerhard Hoffmann\. 1978\.*Gerhard Hoffmann: Raum, Situation, erzählte Wirklichkeit\. Poetologische und historische Studien zum englischen und amerikanischen Roman\.*J\. B\. Metzler\.
- Hripcsak and Rothschild \(2005\)George Hripcsak and Adam S\. Rothschild\. 2005\.[Agreement, the F\-measure, and reliability in information retrieval](https://doi.org/10.1197/jamia.M1733)\.*Journal of the American Medical Informatics Association*, 12\(3\):296–298\.
- Ishikawa and Yoshino \(2025\)Shin\-nosuke Ishikawa and Atsushi Yoshino\. 2025\.[AI with Emotions: Exploring Emotional Expressions in Large Language Models](https://doi.org/10.18653/v1/2025.nlp4dh-1.51)\.In*Proceedings of the 5th International Conference on Natural Language Processing for Digital Humanities*, pages 614–627, Albuquerque, USA\. Association for Computational Linguistics\.
- Ismayilzada et al\. \(2025\)Mete Ismayilzada, Claire Stevenson, and Lonneke van der Plas\. 2025\.[Evaluating Creative Short Story Generation in Humans and Large Language Models](https://doi.org/10.48550/arXiv.2411.02316)\.*arXiv preprint*\.ArXiv:2411\.02316 \[cs\]\.
- Kababgi et al\. \(2024\)Daniel Kababgi, Giulia Grisot, Federico Pennino, and J\. Berenike Herrmann\. 2024\.Recognising nonnamed spatial entities in literary texts: A novel spatial entities classifier\.In*Proceedings of the Computational Humanities Research Conference \(CHR 2024\)*, pages 472–481\.
- Kubinec \(2023\)Robert Kubinec\. 2023\.Ordered beta regression: A parsimonious, well\-fitting model for continuous data with lower and upper bounds\.*Political Analysis*, 31\(4\):519–536\.
- Landis and Koch \(1977\)J\. Richard Landis and Gary G\. Koch\. 1977\.[The measurement of observer agreement for categorical data](https://doi.org/10.2307/2529310)\.*Biometrics*, 33\(1\):159–174\.
- Lee and Lim \(2024\)Bruce W\. Lee and JaeHyuk Lim\. 2024\.[Language Models Don’t Learn the Physical Manifestation of Language](https://doi.org/10.48550/arXiv.2402.11349)\.*arXiv preprint*\.ArXiv:2402\.11349 \[cs\]\.
- Lenth and Piaskowski \(2026\)Russell V\. Lenth and Julia Piaskowski\. 2026\.[*emmeans: Estimated Marginal Means, aka Least\-Squares Means*](https://rvlenth.github.io/emmeans/)\.R package version 2\.0\.3\.
- Lucy and Bamman \(2021\)Li Lucy and David Bamman\. 2021\.[Gender and Representation Bias in GPT\-3 Generated Stories](https://doi.org/10.18653/v1/2021.nuse-1.5)\.In*Proceedings of the Third Workshop on Narrative Understanding*, pages 48–55, Virtual\. Association for Computational Linguistics\.
- Mahlberg \(2013\)Michaela Mahlberg\. 2013\.[*Corpus stylistics and Dickens’s fiction*](https://doi.org/10.4324/9780203076088)\.Routledge advances in corpus linguistics; 14\. Routledge, New York\.
- Marco et al\. \(2024\)Guillermo Marco, Julio Gonzalo, M\.Teresa Mateo\-Girona, and Ramón Del Castillo Santos\. 2024\.[Pron vs Prompt: Can Large Language Models already Challenge a World\-Class Fiction Author at Creative Text Writing?](https://doi.org/10.18653/v1/2024.emnlp-main.1096)In*Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing*, pages 19654–19670, Miami, Florida, USA\. Association for Computational Linguistics\.
- Norris \(1899\)Frank Norris\. 1899\.[*McTeague: A story of San Francisco*](https://www.gutenberg.org/ebooks/165)\.Project Gutenberg\.
- Paech \(2023\)Samuel J\. Paech\. 2023\.[Eq\-bench: An emotional intelligence benchmark for large language models](https://arxiv.org/abs/2312.06281)\.*Preprint*, arXiv:2312\.06281\.
- Rohrbacher \(2025a\)Katrin Rohrbacher\. 2025a\.[de\-corp: A corpus of german\-language fiction and non\-fiction \(1780–1930\)](https://doi.org/10.5334/johd.350)\.*Journal of Open Humanities Data*, 11\(1\):51\.
- Rohrbacher \(2025b\)Katrin Rohrbacher\. 2025b\.[Opening worlds: Narrative beginnings and the role of setting](https://doi.org/10.26083/tuprints-00030149)\.*CCLS2025 Conference Preprints*, 4\(1\)\.
- Rohrbacher \(forthcoming\)Katrin Rohrbacher\. forthcoming\.“lived space”: A computational study of setting in fiction\.In R\. M\. Aust, G\. Grisot, and B\. Herrmann, editors,*Comparing landscapes: Approaches to space and affect in literary fiction*\. Bielefeld University Press\.
- Sears and Weisberg \(2026\)Sydney Sears and Deena Skolnick Weisberg\. 2026\.[Bot or not: Can people tell the difference between stories written by a human or by an AI system?](https://doi.org/10.1017/jdm.2026.10042)*Judgment and Decision Making*, 21:e21\.
- Soni et al\. \(2023\)Sandeep Soni, Amanpreet Sihra, Elizabeth Evans, Matthew Wilkens, and David Bamman\. 2023\.[Grounding characters and places in narrative text](https://doi.org/10.18653/v1/2023.acl-long.655)\.In*Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 11723–11736, Toronto, Canada\. Association for Computational Linguistics\.
- Ströker \(1965\)E\. Ströker\. 1965\.[*Philosophische Untersuchungen zum Raum*](https://books.google.de/books?id=skEAAAAAMAAJ)\.Philosophische Abhandlungen\. V\. Klostermann\.
- Thierry \(1918\)James Francis Thierry\. 1918\.[*The Adventure of the Eleven Cuff\-Buttons*](https://www.gutenberg.org/ebooks/31135)\.Project Gutenberg\.
- Tian et al\. \(2024\)Yufei Tian, Tenghao Huang, Miri Liu, Derek Jiang, Alexander Spangher, Muhao Chen, Jonathan May, and Nanyun Peng\. 2024\.[Are Large Language Models Capable of Generating Human\-Level Narratives?](https://doi.org/10.48550/arXiv.2407.13248)*arXiv preprint*\.ArXiv:2407\.13248 \[cs\]\.
- Underwood et al\. \(2025\)Ted Underwood, Laura K\. Nelson, and Matthew Wilkens\. 2025\.[Can Language Models Represent the Past without Anachronism?](https://doi.org/10.48550/arXiv.2505.00030)*arXiv preprint*\.ArXiv:2505\.00030 \[cs\]\.
- Vaswani et al\. \(2023\)Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N\. Gomez, Lukasz Kaiser, and Illia Polosukhin\. 2023\.[Attention Is All You Need](https://doi.org/10.48550/arXiv.1706.03762)\.*arXiv preprint*\.ArXiv:1706\.03762 \[cs\]\.
- Vauth et al\. \(2021\)Michael Vauth, Hans Ole Hatzel, Evelyn Gius, and Chris Biemann\. 2021\.[Automated event annotation in literary texts](https://ceur-ws.org/Vol-2989/short_paper18.pdf)\.In*Proceedings of the Conference on Computational Humanities Research 2021*, volume 2989 of*CEUR Workshop Proceedings*, pages 333–345, Amsterdam, the Netherlands\. CEUR\-WS\.org\.
- Wang et al\. \(2025\)Qianyue Wang, Jinwu Hu, Zhengping Li, Yufeng Wang, Daiyuan Li, Yu Hu, and Mingkui Tan\. 2025\.[Generating long\-form story using dynamic hierarchical outlining with memory\-enhancement](https://doi.org/10.18653/v1/2025.naacl-long.63)\.In*Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\)*, pages 1352–1391, Albuquerque, New Mexico\. Association for Computational Linguistics\.
- Wankmüller \(2024\)Sandra Wankmüller\. 2024\.[Introduction to Neural Transfer Learning With Transformers for Social Science Text Analysis](https://doi.org/10.1177/00491241221134527)\.*Sociological Methods & Research*, 53\(4\):1676–1752\.Publisher: SAGE Publications Inc\.
- Zhao et al\. \(2024\)Yunpu Zhao, Rui Zhang, Wenyi Li, Di Huang, Jiaming Guo, Shaohui Peng, Yifan Hao, Yuanbo Wen, Xing Hu, Zidong Du, Qi Guo, Ling Li, and Yunji Chen\. 2024\.[Assessing and Understanding Creativity in Large Language Models](https://doi.org/10.1007/s11633-025-1546-4)\.ArXiv:2401\.12491 \[cs\]\.
- Zhong et al\. \(2024\)Shu Zhong, Elia Gatti, Youngjun Cho, and Marianna Obrist\. 2024\.[Exploring Human\-AI Perception Alignment in Sensory Experiences: Do LLMs Understand Textile Hand?](https://doi.org/10.48550/arXiv.2406.06587)*arXiv preprint*\.ArXiv:2406\.06587 \[cs\]\.

## Appendix AEnglish setting classifier

### A\.1Annotation Guidelines

### General Principles

Each sentence is assigned to exactly one category\. Where multiple space types are present, assign the category that is most prominent\. Sentences where space is implied but no concrete or atmospheric markers are present \(e\.g\., references to imagined, dreamt, or remembered spaces\) are labeledno space\. Only spaces that are part of the concrete storyworld are considered\.

### Category Definitions

##### Action space \(Aktionsraum\)

Space that is moved through or interacted with by a character\. Objects serve functional roles, enabling or hindering movement and goal\-directed action\. The character appropriates space through touch and bodily movement rather than through observation or sensory perception\.Example:“He jumped up, jerked the window\-shade, and dragged his chair closer to examine the shoes\.”

##### Perceived space \(gestimmter Raum\)

Space experienced as atmospheric and mood\-laden\. The character is affected by or absorbed into the environment through diffuse sensory experience — sounds, smells, light, weather — without clear directionality or goal\-oriented movement\. Setting may function quasi\-anthropomorphically, as though it acts upon the character\.Example:“The terror of loneliness among those overhanging mountains gripped at the boy’s throat\.”

##### Visual space \(Anschauungsraum\)

Space observed from a static or near\-static position\. The character surveys the environment with their eyes rather than moving through it\. Space presents itself to the character; the focus is on what is seen rather than on how the character is affected\.Example:“He glanced from Tom to the cabin\.”

##### Descriptive space

Spatial information that situates characters or objects without being anchored to any character’s perception, agency, or emotional experience\. Functions as neutral scene\-setting or localization\.Example:“On either side of the towpath were farms and gardens\.”

##### No space

No spatial relationship is present, or space appears only as imagined, dreamt, remembered, or planned, i\.e\., not part of the concrete storyworld\.Example:“By this curious turn I have gained the reputation of deliberate heartlessness\.”

### Common Ambiguities

Action vs\. perceived space\.When movement and atmosphere co\-occur, assign the category that predominates\. If a character moves through space but is primarily affected by or absorbed into it, prefer perceived space\. If movement and goal\-directed interaction with objects are foregrounded, prefer action space\.Perceived vs\. visual space\.Both involve a relatively static character, but visual space is detached and observational, while perceived space involves affective absorption\. If the environment is rendered anthropomorphic or the character is emotionally moved by it, prefer perceived space\.Descriptive vs\. visual space\.Descriptive space is not anchored to any character’s point of view; visual space is\. If a character is explicitly observing the described scene, prefer visual space\.

### A\.2Model and Training

The English setting classifier was fine\-tuned from RoBERTa \(“roberta\-base”\)101010[https://huggingface\.co/FacebookAI/roberta\-base](https://huggingface.co/FacebookAI/roberta-base), an encoder model based on the BERT architecture\([Devlin et al\., 2019](https://arxiv.org/html/2609.02482#bib.bib14)\), trained for 5 epochs with a learning rate of 1e\-5 on 70% of the annotated data and evaluated on a held\-out test set of 30% \(n=972n=972\)\. Classification performance is reported in Table[5](https://arxiv.org/html/2609.02482#A1.T5)\. For comparison, Table[6](https://arxiv.org/html/2609.02482#A1.T6)reports the performance of the German classifier, adapted from[Rohrbacher \(2025b\)](https://arxiv.org/html/2609.02482#bib.bib39), which achieves a macro F1 of 0\.85, broadly similar to the English classifier, with the German classifier performing slightly higher overall\.

Table 5:Classification report for the English setting classifier\. Precision, Recall, and F1\-score are reported per class, along with macro averages\.Table 6:Classification report for the German setting classifier\. Precision, Recall, and F1\-score are reported per class, along with macro averages\. Adapted from[Rohrbacher \(2025b\)](https://arxiv.org/html/2609.02482#bib.bib39)\.
### A\.3Classifier Validation

Since the classifier was fine\-tuned on human\-authored fictional prose, a key question is whether it generalizes to AI\-generated text, which may differ systematically from its training domain in vocabulary or stylistic conventions\. To assess this, two annotators manually labeled a stratified sample of 600 sentences drawn from the AI\-generated texts used in the main study, covering all four models and both languages\. Inter\-annotator agreement was substantial \(κ=0\.736\\kappa=0\.736, 79\.2% raw agreement\), and remaining disagreements were resolved through discussion, producing a gold standard against which classifier performance was evaluated\.

Table[7](https://arxiv.org/html/2609.02482#A1.T7)reports per\-class precision, recall, and F1 for the classifier evaluated against the gold standard annotations\. Table[8](https://arxiv.org/html/2609.02482#A1.T8)shows performance broken down by model and language\. Results are consistent across most model\-language combinations, where accuracy ranges from 77\.3% to 85\.3% withκ\\kappavalues indicating substantial agreement\. This confirms that the classifier generalizes well to AI\-generated text\. Mistral 3\.2 shows somewhat lower performance in both languages, most pronounced in German \(62\.7%,κ=0\.533\\kappa=0\.533\)\. During annotation, we observed that Mistral 3\.2 occasionally produced grammatically malformed or semantically incoherent sentences that superficially resembled a spatial category but could not be assigned one meaningfully\. These were labeled asno spaceby annotators\. The classifier, trained on well\-formed prose, was not exposed to such cases during fine\-tuning and therefore tends to assign a spatial label rather thanno spacein these instances, which disproportionately affects Mistral 3\.2 across both languages and is most pronounced in German where such outputs were most frequent\. Performance for all other models and languages falls within a narrow and acceptable range, supporting the classifier’s general applicability to AI\-generated text\.

Table 7:Per\-class classifier performance against gold standard annotations \(n=600\)\.Table 8:Classifier performance by model and language\.#### A\.3\.1Inter\-annotator agreement

Table 9:Inter\-annotator agreement per category on the double\-coded sample \(n=600n=600sentences\)\. IoU is the number of sentences both annotators assigned to a category divided by the numbern∪n\_\{\\cup\}that at least one annotator did;κ\\kappais the chance\-corrected one\-vs\-rest Cohen’sκ\\kappa\. Then∪n\_\{\\cup\}column sums to 725 rather than 600 because each of the 125 disagreed sentences counts toward two categories\. The underlying marginals are given in Figure[5](https://arxiv.org/html/2609.02482#A1.F5)[5\(a\)](https://arxiv.org/html/2609.02482#A1.F5.sf1)\.Table[9](https://arxiv.org/html/2609.02482#A1.T9)reports two per\-category agreement statistics on the pooled English and German sample\. IoU is a stricter variant of the standard positive\-specific agreement statistic \(inter\-annotator F1;[Cicchetti and Feinstein, 1990](https://arxiv.org/html/2609.02482#bib.bib11);[Hripcsak and Rothschild, 2005](https://arxiv.org/html/2609.02482#bib.bib25)\), which ranks the categories identically\. Perceived space shows the lowest raw agreement \(IoU==0\.543\) but a chance\-correctedκ\\kappaof 0\.646, conventionally substantial agreement\([Landis and Koch, 1977](https://arxiv.org/html/2609.02482#bib.bib30)\)\. These are initial assessment values\. All disagreements were subsequently resolved through discussion, and the classifier was evaluated against the resulting adjudicated gold standard, on which its perceived\-space F1 \(0\.780\) is comparable to the other categories \(0\.720–0\.826; Table[7](https://arxiv.org/html/2609.02482#A1.T7)\)\.

![Refer to caption](https://arxiv.org/html/2609.02482v1/figure_confusion_matrix_iaa.png)\(a\)Annotator 1 \(rows\) against Annotator 2 \(columns\) on the double\-coded sample,n=600n=600\.
![Refer to caption](https://arxiv.org/html/2609.02482v1/figure_confusion_matrix_by_language.png)\(b\)Classifier predictions \(columns\) against the adjudicated gold standard \(rows\),n=300n=300per language\.

Figure 5:Confusion structure of the five\-category scheme\. Panel[5\(a\)](https://arxiv.org/html/2609.02482#A1.F5.sf1)reports human–human agreement; panel[5\(b\)](https://arxiv.org/html/2609.02482#A1.F5.sf2)reports classifier–gold agreement\.Figure[5](https://arxiv.org/html/2609.02482#A1.F5)shows the confusion structure behind these agreement figures, which suggests that the disagreements follow a systematic pattern\. In panel[5\(a\)](https://arxiv.org/html/2609.02482#A1.F5.sf1), perceived space is confused principally with no space \(21 sentences\) and with visual space \(18\)\. Both confusions involve threshold judgments, one about whether atmospheric content is concretely spatial at all, the other about affective absorption as against detached observation \(Appendix[A\.1](https://arxiv.org/html/2609.02482#A1.SS1)\)\. Perceived and action space, the two categories carrying the central contrast of this paper, are almost never confused\. The count is 8 of 600 sentences \(1\.3%\) between annotators, and 4 of 300 \(English\) and 2 of 300 \(German\) for the classifier against the gold standard\. The reported difference is therefore unlikely to be an artefact of annotator or classifier uncertainty between these two categories\. Panel[5\(b\)](https://arxiv.org/html/2609.02482#A1.F5.sf2)shows that the classifier’s errors are concentrated in the no space row\. Recall for that category is 0\.611 in English \(55 of 90\) and 0\.564 in German \(57 of 101\)\. Misassigned sentences go most often to perceived space, and over half of these come from Mistral 3\.2, whose outputs are frequently malformed \(Section[A\.3](https://arxiv.org/html/2609.02482#A1.SS3)\)\. Only four occur across all 150 GPT 4\.1 sentences\. Because gold no\-space sentences are by construction those with the weakest spatial signal, this asymmetry inflates the measured amount of represented space, but is concentrated in Mistral 3\.2, whose deviations are among the smallest of the four models reported in Section[4\.3](https://arxiv.org/html/2609.02482#S4.SS3)\. The central contrast is unaffected\.

## Appendix BDeviation from Human Baseline: English and German

Figure[6](https://arxiv.org/html/2609.02482#A2.F6)shows the deviation of each model from the human baseline across ten narrative sections for English and German\.

Figure 6:Deviation of AI model predictions from the human baseline across narrative sections for each spatial category, for English \(top\) and German \(bottom\)\. The human baseline represents the average normalized frequency of each space type per narrative section \(deviation = 0\); positive values indicate higher proportions than human\-authored texts, negative values indicate lower\. Error bars indicate 95% confidence intervals\.
## Appendix CPrompt Sensitivity

To validate that the findings are not sensitive to prompt formulation, we ran three prompt variants \(see Table[10](https://arxiv.org/html/2609.02482#A3.T10)\) per open\-source model and language in addition to the original prompt\. Figures[7](https://arxiv.org/html/2609.02482#A3.F7)and[8](https://arxiv.org/html/2609.02482#A3.F8)show the normalized frequency of each spatial category across ten narrative sections for the original prompt and three variants, for English and German respectively\. Although prompt formulation introduces some variation in the absolute frequency of each category, the overall trends across the narrative sections remain consistent across all conditions\. This suggests that the patterns reported in the main analysis are not an artifact of the specific prompt used\.

Figure 7:Normalized frequency of spatial categories across ten narrative sections \(English\)\. Human\-authored fiction is shown as a single line; for each model, four lines represent the original prompt and three prompt variants\.Figure 8:Normalized frequency of spatial categories across ten narrative sections \(German\)\. Human\-authored fiction is shown as a single line\. For each model, four lines represent the original prompt and three prompt variants\.Table 10:Prompt variants used in the prompt sensitivity ablation\. V1–V3 vary in phrasing and verbosity\.\{genre\}and\{sentence\}are filled with metadata and the opening sentence of each source text\. Step 2 is repeated for each subsequent chapter\.
## Appendix DTemporal Distribution of Setting across Narrative Time

Figure[9](https://arxiv.org/html/2609.02482#A4.F9)shows the full temporal distribution of setting categories across narrative time for English and German\.

![Refer to caption](https://arxiv.org/html/2609.02482v1/figures/action_space_eng.png)
![Refer to caption](https://arxiv.org/html/2609.02482v1/figures/perceived_space_eng.png)
![Refer to caption](https://arxiv.org/html/2609.02482v1/figures/descriptive_space_eng.png)
![Refer to caption](https://arxiv.org/html/2609.02482v1/figures/visual_space_eng.png)
![Refer to caption](https://arxiv.org/html/2609.02482v1/figures/all_space_eng.png)

![Refer to caption](https://arxiv.org/html/2609.02482v1/figures/action_space_germ.png)
![Refer to caption](https://arxiv.org/html/2609.02482v1/figures/perceived_space_germ.png)
![Refer to caption](https://arxiv.org/html/2609.02482v1/figures/descriptive_space_germ.png)
![Refer to caption](https://arxiv.org/html/2609.02482v1/figures/visual_space_germ.png)
![Refer to caption](https://arxiv.org/html/2609.02482v1/figures/all_space_germ.png)

Figure 9:Temporal distribution of setting categories across narrative time for English \(left\) and German \(right\)\.
## Appendix EModel Identifiability from Spatial Distributions

A random forest classifier trained on per\-section spatial category proportions identified the source model well above chance in both languages \(confusion matrices in Figures[10\(a\)](https://arxiv.org/html/2609.02482#A5.F10.sf1)and[10\(b\)](https://arxiv.org/html/2609.02482#A5.F10.sf2)\)\. In English, human\-authored texts are most reliably identified \(recall = 0\.97\), while Mistral 3\.2 is hardest \(0\.56\)\. The main source of confusion is between LlaMA 3\.3 and Mistral 3\.2, consistent with their closely aligned setting distributions\. In German, human recall drops to 0\.79 as LLM distributions converge toward the baseline\.

![Refer to caption](https://arxiv.org/html/2609.02482v1/confusion_matrix_en.png)\(a\)English
![Refer to caption](https://arxiv.org/html/2609.02482v1/confusion_matrix_ger.png)\(b\)German

Figure 10:Row\-normalized confusion matrices for model identification from spatial category distributions\.
## Appendix FTemporal Stability of Human\-Authored Corpora

Figure[11](https://arxiv.org/html/2609.02482#A6.F11)shows the distribution of setting categories across publication years in the human\-authored corpora for English and German\. Distributions are broadly stable across the main period of each corpus, supporting the assumption that observed differences between human\-authored and AI\-generated texts are not driven by historical sub\-period\.

\(a\)English corpus \(1800–1920\)\.\(b\)German corpus \(1780–1940\), reproduced from[Rohrbacher \(forthcoming\)](https://arxiv.org/html/2609.02482#bib.bib40)\.
Figure 11:Normalized frequencies of setting categories across publication years in the human\-authored corpora\. Shaded bands indicate 95% confidence intervals\. Wider intervals in earlier decades reflect smaller sample sizes\. Distributions are broadly stable across the main period of each corpus\.
## Appendix GFull GLMM Results

Table[11](https://arxiv.org/html/2609.02482#A7.T11)reports the full Type III Waldχ2\\chi^\{2\}tests for the GLMM described in Section[3\.5](https://arxiv.org/html/2609.02482#S3.SS5), fitted separately for each setting category and language\. These are the models from which all test statistics reported in Section[4\.3](https://arxiv.org/html/2609.02482#S4.SS3)are taken\. Genre is coded as novel vs\. other fiction and, as noted in Section[3\.5](https://arxiv.org/html/2609.02482#S3.SS5), is matched across author conditions by construction\.

Table 11:Type III Waldχ2\\chi^\{2\}tests for the GLMM,space ~ author×\\timessection \+ genre \+ \(1 \| story\), fitted separately per setting category and language\.∗∗∗p<\.001\{\}^\{\*\*\*\}p<\.001,∗p<\.05\{\}^\{\*\}p<\.05, n\.s\.==not significant\.
## Appendix HChapter\-Boundary Position

Figures[12](https://arxiv.org/html/2609.02482#A8.F12)and[13](https://arxiv.org/html/2609.02482#A8.F13)show the profiles for perceived space by position within chapter, for English and German respectively, and Table[12](https://arxiv.org/html/2609.02482#A8.T12)reports the corresponding means\.

Beyond perceived space, action space shows the mirror pattern, dropping at chapter boundaries relative to chapter middles, while visual and descriptive space decline slightly and steadily across each chapter, with a total spread of at most 0\.006 and no boundary peak\. The boundary effect is thus specific to perceived and, inversely, action space\.

Figure 12:Normalized frequency of perceived space by position within chapter, English\. Each chapter is divided into four equal\-length sections and the four chapters are shown consecutively\. The dashed line indicates the human baseline\.Figure 13:Normalized frequency of perceived space by position within chapter, German\. Each chapter is divided into four equal\-length sections and the four chapters are shown consecutively\. The dashed line indicates the human baseline\.Table 12:Normalized frequency of perceived space by quarter within chapter \(Q1–Q4\), averaged over the four chapters\. Human\-authored texts are not chapter\-segmented; their corpus mean is given as a single reference value\.
## Appendix IGeneration Length

We generated between 2,500 and 3,000 tokens per chapter with an enforced maximum of 3,000 tokens, stopping when the model emitted an end\-of\-sequence token or reached the cap, for four chapters and a total of 10,000–12,000 tokens per story\. Where a model reproduced the seeding sentence verbatim, we removed the duplicated first sentence; where generation stopped mid\-sentence, we removed the trailing incomplete sentence\. Table[13](https://arxiv.org/html/2609.02482#A9.T13)reports the resulting story lengths in words\.

Table 13:Story length in words per model and language \(n=1,000n=1\{,\}000each\)\.

Similar Articles

Do Large Language Models Always Tell The Same Stories?

arXiv cs.CL

This paper investigates whether large language models generate diverse stories. Using narrative similarity analysis, the authors find that LLM-generated narratives are consistently more similar to each other than human-written stories, and that common mitigation strategies like negative prompting and temperature scaling fail to address this homogeneity.