Tag
This paper introduces a new multimodal benchmark for evaluating AI models' narrative understanding of Hollywood films, addressing copyright issues, and shows that vision-language and audio-visual models perform below human-level accuracy on the task.
NARU is a benchmark for evaluating narrative evolution and cultural reasoning in Japanese extreme long-form videos, constructed through a hierarchical annotation pipeline and native-speaker verification to address gaps in current benchmarks.
SwanNLP presents an LLM-based framework for plausibility scoring in narrative word sense disambiguation at SemEval-2026 Task 5, using structured reasoning and dynamic few-shot prompting to predict human-perceived plausibility of word senses in short stories. The work demonstrates that commercial large-parameter LLMs with few-shot prompting and model ensembling effectively replicate human judgment patterns in realistic narrative contexts.