ShotPlan: Cinematic Video Generation with Learnable Planning Token
Summary
ShotPlan introduces learnable planning tokens with fractional temporal rotary position embeddings for cinematic multi-shot video generation, enabling explicit shot-level planning and achieving superior inter-shot consistency.
View Cached Full Text
Cached at: 07/21/26, 02:37 PM
Paper page - ShotPlan: Cinematic Video Generation with Learnable Planning Token
Source: https://huggingface.co/papers/2607.17675 Published on Jul 20
·
Submitted byhttps://huggingface.co/Pensioner
suguoon Jul 21
Abstract
Currentvideogenerationmodelsachieveimpressiveresultsinsingle-shotgeneration,yetremainlimitedincinematicvideogeneration,wherecoherentnarrativesandeffectivemulti-shotcompositionrequireexplicitshotplanning.Toaddressthischallenge,weproposeShotPlan,aframeworkforexplicitmulti-shotcinematicvideogenerationbuiltuponavideodiffusionfoundationmodel.Ourmethodintroduceslearnableplanningtokensthatcaptureshot-leveltransitioncuesandcanbeseamlesslyintegratedwiththeoriginalvideogenerationtokenstocontroltransitiontimestamps.Unlikestandardvideogenerationtokens,theproposedplanningtokensareequippedwithFractionalTemporalRotaryPositionEmbedding(FRoPE),enablingshottransitionstobemodeledattheframelevel.ExperimentsdemonstratethatShotPlansignificantlyoutperformsexistingcinematicvideogenerationmethods,offeringmoreflexibleshotmanagementandstrongerinter-shotconsistency.
View arXiv pageView PDFProject pageGitHub5Add to collection
Get this paper in your agent:
hf papers read 2607\.17675
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.17675 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.17675 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.17675 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Experimenting with storyboard-planned AI cinematics instead of single-prompt generation
Explores a storyboard-planned approach for AI cinematics that builds sequence structure before generating shots individually, resulting in more coherent video compared to single-prompt generation, while noting current weaknesses like identity drift and interaction physics.
SmartDirector: Keyframe-Conditioned Cinematic Video Generation with Narrative Pacing Control
SmartDirector is a framework that enhances video generation by using multiple keyframes to improve narrative structure and temporal pacing, operating in a two-stage process of low-resolution generation and high-resolution refinement.
CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives
CausalCine is a new academic framework for real-time, interactive multi-shot video generation that uses causal modeling and dynamic memory routing to improve cross-shot coherence in autoregressive models.
Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning
This paper introduces an agentic framework combining LLMs and VLMs for consistent multi-instruction video editing across multiple shots, and proposes the MMLVE task and benchmark to evaluate performance.
@heyshrutimishra: Start with AI Video. Build the next scene around the story you want to tell. Not just another visually interesting shot…
Seedance 2.5 is highlighted for AI video generation, allowing precise placement of actions, reveals, and transitions via timestamp instructions to support story-driven scenes.