Tag
This paper introduces WSE-bench, a process benchmark for evaluating LLMs in open-ended world simulations, separately assessing sustained generation, canonical coherence, and meaningful development.
This paper introduces ArcANE, an automatically constructed benchmark for evaluating role-playing language agents' alignment with character psychological trajectories across narrative phases, showing that conditioning on character arc information improves performance, especially in scenarios beyond the source text.
A preprint proposes a 33-feature quantitative linguistic framework that distinguishes professionally edited from self-published books and outperforms existing story-level evaluation metrics.