ConWriter: Transition-Constrained Stateful Long-Form Story Generation with Lightweight Neuro-Symbolic Consistency Control
Summary
ConWriter introduces a training-free framework for long-form story generation that maintains narrative consistency through scene-level incremental writing, symbolic state reasoning, and uncertainty-aware risk signals. Evaluated on ConStory-Bench across multiple models and lengths, it aims to prevent consistency errors from propagating in extended contexts.
View Cached Full Text
Cached at: 08/07/26, 07:49 AM
# ConWriter: Transition-Constrained Stateful Long-Form Story Generation with Lightweight Neuro-Symbolic Consistency Control Source: [https://arxiv.org/abs/2608.05169](https://arxiv.org/abs/2608.05169) [View PDF](https://arxiv.org/pdf/2608.05169) > Abstract:Long\-form story generation requires models to preserve narrative consistency across extended contexts, yet existing prompting\-based methods often accumulate temporal, factual, character, commonsense, and stylistic errors as the story grows\. We propose ConWriter, a training\-free framework for consistency\-aware long story generation\. ConWriter writes stories incrementally at the scene level, guided by static story requirements, dynamic narrative memory, symbolic state reasoning, and uncertainty\-aware risk signals\. Rather than treating long\-story generation as a single free\-form decoding process, ConWriter maintains evolving story states, checks whether new scenes satisfy required narrative transitions, and uses uncertainty\-aware risk signals to prioritize validation and localized repair\. This enables consistency control during generation, before local errors propagate into later scenes\. We evaluate ConWriter on ConStory\-Bench, covering four long\-story tasks: continuation, generation, expansion, and completion\. Due to the high cost of long\-form generation and evaluation, we use the first five cases from each task and test 3k, 6k, and 12k target lengths across Qwen3\.5\-Plus, DeepSeek\-V4\-Flash, and GPT\-5 series\. Experiments follow the official ConStory\-Bench evaluation protocol\. ## Submission history From: Jindong Li \[[view email](https://arxiv.org/show-email/8e1baefc/2608.05169)\] **\[v1\]**Wed, 27 May 2026 17:53:18 UTC \(3,295 KB\)
Similar Articles
World-State Transformations for Neuro-symbolic Interactive Storytelling
This paper explores using LLMs to predict state changes within rule-based interactive storytelling systems, aiming to improve coherence and player expression. Experiments with Llama 3 70B and Gemini 1.5 Flash show that world-state transformations can maintain consistency while encouraging creative player input.
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives
This paper introduces NCP-Bench, a benchmark derived from 100 movie synopses for evaluating long-horizon narrative consistency in LLM-based interactive storytelling agents. Experiments show that even strong models like GPT-5.2 struggle to maintain logical consistency, with a 42% survival rate after 20 turns and high fact conflict rates.
Narrative World Model: Narratology-Grounded Writer Memory for Long-Form Fiction
This paper introduces the Narrative World Model (NWM), a memory system for long-form fiction writers that uses narratology-grounded typed temporal-state graphs and query-conditioned hybrid retrieval to answer multi-hop questions about evolving story state. The system significantly outperforms existing temporal-knowledge-graph frameworks like Graphiti on benchmark narratological QA tasks.
Towards Human-Level Book-Writing Capability
This paper introduces a dataset and training framework that transforms human-authored novels into multi-resolution planning scaffolds, enabling long-context language models to generate book-scale fiction with more human-like prose and narrative dynamics.
Controllable Narrative Rendering for Enhanced Assisted Writing
This paper introduces Loom, an assisted writing framework that leverages a three-layer pipeline based on narratological story/discourse distinction to control narrative intent and rendering density, achieving improved factual integrity and descriptive intensity compared to baselines.