ThinkV2V: Unleashing the Reasoning Capability of MLLMs for Instruction-Guided Video Editing

Hugging Face Daily Papers Papers

Summary

ThinkV2V introduces a reasoning-driven framework that activates MLLM thinking before visual generation for instruction-guided video editing, using an MLLM-to-DiT architecture with progressive curriculum training and inference-time thinking scaling. The authors also release the ThinkV2V-150K dataset and ThinkV2V-Bench, showing their 5B DiT model outperforms larger 10B baselines.

Instruction-guided video editing has made significant progress, yet existing methods use multimodal large language models (MLLMs) primarily as semantic encoders, so they often fall short in working with implicit edits that require causal or semantic reasoning. To bridge this fundamental gap in video editing, we propose ThinkV2V, a reasoning-driven framework for complex instruction-guided video editing, explicitly activating MLLM thinking before visual generation. At its core, ThinkV2V builds on a practical MLLM-to-DiT architecture to turn explicit thinking over the source video and instruction into refined conditioning signals for video editing. Further, we equip it with a dedicated training and inference recipe, combining Progressive Curriculum Training, which gradually cultivates the model from basic editing to reasoning-intensive cases, with Inference-Time Thinking Scaling, which iteratively refines candidate prompts and selects the most reliable one, to better elicit reasoning in challenging editing scenarios. We also curate the ThinkV2V-150K dataset and introduce ThinkV2V-Bench to support training and evaluation of video editing with implicit intent and causal reasoning. Experimental results demonstrate the state-of-the-art performance of ThinkV2V on both complex and standard editing scenarios, in which our 5B-scale DiT model substantially outperforms larger 10B-scale baselines.
Original Article
View Cached Full Text

Cached at: 10/01/26, 04:21 AM

Paper page - ThinkV2V: Unleashing the Reasoning Capability of MLLMs for Instruction-Guided Video Editing

Source: https://huggingface.co/papers/2609.38541 Authors:

,

,

,

,

,

,

,

,

,

,

,

Abstract

Instruction-guidedvideoeditinghasmadesignificantprogress,yetexistingmethodsusemultimodallargelanguagemodels(MLLMs)primarilyassemanticencoders,sotheyoftenfallshortinworkingwithimpliciteditsthatrequirecausalorsemanticreasoning.Tobridgethisfundamentalgapinvideoediting,weproposeThinkV2V,areasoning-drivenframeworkforcomplexinstruction-guidedvideoediting,explicitlyactivatingMLLMthinkingbeforevisualgeneration.Atitscore,ThinkV2VbuildsonapracticalMLLM-to-DiTarchitecturetoturnexplicitthinkingoverthesourcevideoandinstructionintorefinedconditioningsignalsforvideoediting.Further,weequipitwithadedicatedtrainingandinferencerecipe,combiningProgressiveCurriculumTraining,whichgraduallycultivatesthemodelfrombasiceditingtoreasoning-intensivecases,withInference-TimeThinkingScaling,whichiterativelyrefinescandidatepromptsandselectsthemostreliableone,tobetterelicitreasoninginchallengingeditingscenarios.WealsocuratetheThinkV2V-150KdatasetandintroduceThinkV2V-Benchtosupporttrainingandevaluationofvideoeditingwithimplicitintentandcausalreasoning.Experimentalresultsdemonstratethestate-of-the-artperformanceofThinkV2Vonbothcomplexandstandardeditingscenarios,inwhichour5B-scaleDiTmodelsubstantiallyoutperformslarger10B-scalebaselines.

View arXiv pageView PDFProject pageGitHub4Add to collection

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.38541 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.38541 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.38541 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning

Hugging Face Daily Papers

The paper introduces Internalized Visual Thinking (IVT), a post-training framework that trains multimodal models to predict future frame embeddings, enabling direct answer generation without synthesizing intermediate images, thereby reducing latency over 5x for proactive video reasoning.

A Very Big Video Reasoning Suite

Papers with Code Trending

This paper introduces the Very Big Video Reasoning (VBVR) dataset and benchmark, a large-scale resource with over one million video clips across 200 reasoning tasks, enabling systematic study of spatiotemporal reasoning and showing early signs of emergent generalization.