MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation

Hugging Face Daily Papers 05/09/26, 12:00 AM Papers

Summary

MuSS introduces a large-scale dataset and benchmark for multi-shot subject-to-video generation, addressing narrative logic and copy-paste issues in cinematic storytelling.

While video foundation models excel at single-shot generation, real-world cinematic storytelling inherently relies on complex multi-shot sequencing. Further progress is constrained by the absence of datasets that address three core challenges: authentic narrative logic, spatiotemporal text-video alignment conflicts, and the "copy-paste" dilemma prevalent in Subject-to-Video (S2V) generation. To bridge this gap, we introduce MuSS, a large-scale, dual-track dataset tailored for multi-shot video and S2V generation. Sourced from over 3,000 movies, MuSS explicitly supports both complex montage transitions and subject-centric narratives. To construct this dataset, we pioneer a progressive captioning pipeline that eliminates contextual conflicts by ensuring local shot-level accuracy before enforcing global narrative coherence. Crucially, we implement a cross-shot matching mechanism to fundamentally eradicate the S2V copy-paste shortcut. Alongside the dataset, we propose the Cinematic Narrative Benchmark, featuring a visual-logic-driven paradigm and a novel Anti-Copy-Paste Variance (ACP-Var) metric to rigorously assess continuous storytelling and 3D structural consistency. Extensive experiments demonstrate that while current baselines struggle with continuous narrative logic or degenerate into trivial 2D sticker generators, our MuSS-augmented model achieves state-of-the-art narrative effectiveness and cross-shot identity preservation.

Original Article

View Cached Full Text

Cached at: 05/12/26, 07:30 AM

Paper page - MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation

Source: https://huggingface.co/papers/2604.23789

Abstract

MuSS is a large-scale dual-track dataset designed for multi-shot video generation that addresses narrative logic, spatiotemporal alignment, and copy-paste issues in subject-to-video generation through a progressive captioning pipeline and cross-shot matching mechanism.

While video foundation models excel at single-shot generation, real-world cinematic storytelling inherently relies on complex multi-shot sequencing. Further progress is constrained by the absence of datasets that address three core challenges: authenticnarrative logic,spatiotemporal text-video alignmentconflicts, and the “copy-paste” dilemma prevalent inSubject-to-Video (S2V) generation. To bridge this gap, we introduce MuSS, a large-scale,dual-track datasettailored for multi-shot video and S2V generation. Sourced from over 3,000 movies, MuSS explicitly supports both complex montage transitions and subject-centric narratives. To construct this dataset, we pioneer aprogressive captioning pipelinethat eliminates contextual conflicts by ensuring local shot-level accuracy before enforcing global narrative coherence. Crucially, we implement across-shot matching mechanismto fundamentally eradicate the S2V copy-paste shortcut. Alongside the dataset, we propose theCinematic Narrative Benchmark, featuring a visual-logic-driven paradigm and a novelAnti-Copy-Paste Variance (ACP-Var) metricto rigorously assesscontinuous storytellingand3D structural consistency. Extensive experiments demonstrate that while current baselines struggle with continuousnarrative logicor degenerate into trivial 2D sticker generators, our MuSS-augmented model achieves state-of-the-art narrative effectiveness and cross-shot identity preservation.

View arXiv page View PDF GitHub5 Add to collection

Get this paper in your agent:

hf papers read 2604\.23789

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2604.23789 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2604.23789 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2604.23789 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation

Paper page - MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation

Abstract

Models citing this paper0

Datasets citing this paper0

Spaces citing this paper0

Collections including this paper0

Similar Articles

Memento: Reconstruct to Remember for Consistent Long Video Generation

CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation

Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos

VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis

Submit Feedback

Similar Articles

Memento: Reconstruct to Remember for Consistent Long Video Generation

CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation

Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos

VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis