标签
This paper introduces NCP-Bench, a benchmark derived from 100 movie synopses for evaluating long-horizon narrative consistency in LLM-based interactive storytelling agents. Experiments show that even strong models like GPT-5.2 struggle to maintain logical consistency, with a 42% survival rate after 20 turns and high fact conflict rates.
本文介绍了 NARRA-Gym,这是一个基准和可执行评估环境,用于评估大型语言模型在多轮对话中维持交互式叙事、管理记忆以及适应用户的能力。