NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video
Summary
NARU is a benchmark for evaluating narrative evolution and cultural reasoning in Japanese extreme long-form videos, constructed through a hierarchical annotation pipeline and native-speaker verification to address gaps in current benchmarks.
View Cached Full Text
Cached at: 08/21/26, 12:09 PM
Paper page - NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video
Source: https://huggingface.co/papers/2608.13210
Abstract
NARU is a Japanese long-form video benchmark evaluating narrative evolution and cultural reasoning through a hierarchical annotation pipeline and extensive native-speaker verification.
Long-form video understandingencompasses tasks that go beyond retrieving isolated events, including tracking an evolving narrative and interpreting social meaning that may remain implicit. However, existing benchmarks rarely evaluate these capabilities jointly, particularly in high-context, non-English media. To address this gap, we introduce NARU, a benchmark designed to evaluateNarrative evolutionand Reasoning on cultural Understanding in Japanese long-form video. NARU consists of 1,481 questions grounded in 155 videos totaling 146.8 hours, spanning four narrative and five cultural dimensions. To construct the benchmark at this scale, we propose ahierarchical memory-based annotationpipeline that transforms raw video into structured event, narrative, and cultural annotations, then generates questions viatask-oriented synthesisand iterativeshortcut removal. The construction process includes two native-speaker verification stages involving 68 annotators. Evaluations across eight model configurations reveal substantial limitations in both long-range narrative integration and culturally grounded reasoning. By exposing these persistent gaps, NARU offers a systematic testing ground for developingMLLMscapable of reliably interpreting long-form, high-context video.
View arXiv pageView PDFProject pageGitHub3Add to collection
Get this paper in your agent:
hf papers read 2608\.13210
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.13210 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2608.13210 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.13210 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Characterizing Narrative Content in Web-scale LLM Pretraining Data
A fine-grained study of narrative features in web-scale LLM pretraining data, introducing NarraBERT and NarraDolma to measure narrative patterns and their distribution across sources.
NARRA-Gym for Evaluating Interactive Narrative Agents
This paper introduces NARRA-Gym, a benchmark and executable evaluation environment for assessing Large Language Models' abilities in sustaining interactive narratives, managing memory, and adapting to users over multiple turns.
Narrative Knowledge Weaver: Narrative-Centric Retrieval-Augmented Reasoning for Long-Form Text Understanding
Introduces Narrative Knowledge Weaver (NKW), a source-grounded framework for narrative-centric retrieval-augmented reasoning in long-form text understanding. It aligns textual evidence, atomic facts, graph structure, entity profiles, and storylines, achieving strong results on screenplay-level story-world QA benchmarks.
NarrativeWorldBench: A Frontier-Saturated Benchmark and a Latent World Model for Long-Horizon Co-Creative Audio Drama
This paper introduces NarrativeWorldBench, a benchmark for evaluating long-horizon narrative consistency in audio dramas, and N-VSSM, a latent state-space model that outperforms frontier LLMs across multiple horizons and languages.
Show HN: FastUbu – An Ultrafast Video Archive
FastUbu is a tool that applies modern AI techniques like indexing and transcription to the 30-year-old Ubu film archive, aiming to provide ultrafast video processing via the Kino API.