Tag
ConWriter introduces a training-free framework for long-form story generation that maintains narrative consistency through scene-level incremental writing, symbolic state reasoning, and uncertainty-aware risk signals. Evaluated on ConStory-Bench across multiple models and lengths, it aims to prevent consistency errors from propagating in extended contexts.
Introduces EpiNarrate, an agentic framework that separates structured numerical reasoning from natural-language generation to produce grounded epidemiological narratives from ensemble projections.
This paper introduces the Narrative World Model (NWM), a memory system for long-form fiction writers that uses narratology-grounded typed temporal-state graphs and query-conditioned hybrid retrieval to answer multi-hop questions about evolving story state. The system significantly outperforms existing temporal-knowledge-graph frameworks like Graphiti on benchmark narratological QA tasks.
Introduces Magnet, a multi-agent goal-driven narrative engine for long-form story generation with persona-grounded characters, and Atlas, a graph-based pipeline for detecting hallucinations in generated narratives. The framework improves coherence and reduces hallucinations compared to single-model baselines and IBSEN.
This paper introduces Loom, an assisted writing framework that leverages a three-layer pipeline based on narratological story/discourse distinction to control narrative intent and rendering density, achieving improved factual integrity and descriptive intensity compared to baselines.
This paper introduces Narrative-UFET, a method that generates short narratives to provide broader context for ultra-fine entity typing, improving performance on long-tail types compared to sentence-level baselines.
This paper proposes ReverieMem, a three-layer memory architecture for book-based LLM role-playing agents that prevents factual overreach and stylistic monotony. It also introduces the KBF-QA benchmark and achieves significant improvements in knowledge boundary fidelity and narrative quality.
A hierarchical multi-agent framework generates short dramas from single sentences by enforcing narrative pacing, ensuring spatial consistency, and implementing quality control through iterative refinement and reviewer loops. It introduces a new benchmark, Short-Drama-Bench, for evaluation.
Researchers introduce BIASEDTALES-ML, a large-scale multilingual dataset of ~350,000 LLM-generated children's stories across eight languages, designed to analyze narrative attribute distributions and cross-lingual bias patterns in language model outputs. The work reveals significant cross-lingual variability, highlighting limitations of English-centric bias evaluations.
This paper proposes CAP-TTA, a test-time adaptation framework that uses preconditioned LoRA updates triggered by bias-risk scores to mitigate toxicity and bias in large language models during narrative generation, achieving faster optimization and better fluency than standard baselines.
ArcDeck is a multi-agent framework that generates presentation slides from academic papers by modeling logical flow through discourse trees and iterative agent refinement, outperforming direct summarization methods. The paper introduces ArcBench, a new benchmark for evaluating paper-to-slide generation with emphasis on narrative coherence and logical structure.