Towards a Linguistic Evaluation of Narratives: A Quantitative Stylistic Framework
Summary
A preprint proposes a 33-feature quantitative linguistic framework that distinguishes professionally edited from self-published books and outperforms existing story-level evaluation metrics.
View Cached Full Text
Cached at: 04/22/26, 08:30 AM
# Towards a Linguistic Evaluation of Narratives: A Quantitative Stylistic Framework Source: [https://arxiv.org/abs/2604.19261](https://arxiv.org/abs/2604.19261) [View PDF](https://arxiv.org/pdf/2604.19261) > Abstract:The evaluation of narrative quality remains a complex challenge, as it involves subjective factors such as plot, character development, and emotional impact\. This work proposes a quantitative approach to narrative assessment by focusing on the linguistic dimension as a primary indicator of quality\. The paper presents a methodology for the automatic evaluation of narrative based on the extraction of a comprehensive set of 33 quantitative linguistic features categorized into lexical, syntactic, and semantic groups\. To test the model, an experiment was conducted on a specialized corpus of 23 books, including canonical masterpieces and self\-published works\. Through a similarity matrix, the system successfully clustered the narratives, distinguishing almost perfectly between professionally edited and self\-published texts\. Furthermore, the methodology was validated against a human\-annotated dataset; it significantly outperforms traditional story\-level evaluation metrics, demonstrating the effectiveness of quantitative linguistic features in assessing narrative quality\. ## Submission history From: Alessandro Maisto \[[view email](https://arxiv.org/show-email/2b9a5b8d/2604.19261)\] **\[v1\]**Tue, 21 Apr 2026 09:21:40 UTC \(827 KB\)
Similar Articles
@emollick: There is a lot being written about the stylistic tells of AI writing (em-dashes, etc.) but this paper looks at AI narra…
This paper introduces StoryScope, a pipeline that analyzes discourse-level narrative features to distinguish AI-generated fiction from human-written stories. It achieves high accuracy and reveals distinct narrative fingerprints for different LLMs like Claude, GPT, and Gemini.
Beyond LLMs: A Linguistic Approach to Causal Graph Generation from Narrative Texts
This paper proposes a hybrid framework combining LLM summarization, a linguistically grounded Expert Index, and STAC classification to generate causal graphs from narrative texts, outperforming GPT-4o and Claude 3.5 on benchmark stories.
SAGE: Hierarchical LLM-Based Literary Evaluation through Ontology-Grounded Interpretive Dimensions
This paper introduces SAGE, a hierarchical LLM-based framework for evaluating literary quality through ontology-grounded interpretive dimensions. It demonstrates high reliability and inter-rater agreement in assessing cultural, emotional, and philosophical aspects of narratives, highlighting gaps between human-authored and LLM-generated works.
Comprehensive Evaluation of Large Language Model Responses: A Multi-Factor Scoring System
This paper proposes a multi-factor scoring system for evaluating LLM responses, integrating accuracy, conciseness, factual consistency, readability, and coherence. Applied to the TruthfulQA dataset, it reveals strengths and limitations of mainstream models, offering a transparent evaluation framework.
Beyond Accuracy: A Multidimensional Evaluation of Statistical Reasoning in Large Language Models
This paper proposes a multidimensional evaluation framework for assessing statistical reasoning in large language models, combining response accuracy, response behavior, structural topic modeling, and lexical similarity analysis across 15 LLMs and 90 exam questions. It finds that accuracy alone is insufficient to characterize LLM statistical reasoning and that vendor-specific stylistic differences exist.