Tag
ShotPlan introduces learnable planning tokens with fractional temporal rotary position embeddings for cinematic multi-shot video generation, enabling explicit shot-level planning and achieving superior inter-shot consistency.
This paper shows that the reasoning gap between base LLMs and large reasoning models is concentrated on a small set of early planning tokens. It introduces disagreement-guided token intervention, where replacing only those critical tokens with a reasoning model's outputs allows a base model to nearly match the reasoning model's performance.