🤖 AI Summary
Existing methods typically employ multimodal large language models (MLLMs) solely as semantic encoders, struggling with implicit video editing instructions that require causal or semantic reasoning. This work proposes ThinkV2V, a novel MLLM-to-DiT reasoning framework that pioneers the activation of explicit MLLM thinking to generate refined conditioning signals. By integrating progressive curriculum learning with test-time thought expansion strategies, ThinkV2V enables complex instruction-driven video editing. Additionally, this project releases the ThinkV2V-150K dataset alongside a dedicated evaluation benchmark. Experimental results demonstrate that the proposed approach achieves state-of-the-art performance across both complex and standard scenarios, with a 5B-parameter model significantly outperforming 10B-parameter baselines.
📝 Abstract
Instruction-guided video editing has made significant progress, yet existing methods use multimodal large language models (MLLMs) primarily as semantic encoders, so they often fall short in working with implicit edits that require causal or semantic reasoning. To bridge this fundamental gap in video editing, we propose ThinkV2V, a reasoning-driven framework for complex instruction-guided video editing, explicitly activating MLLM thinking before visual generation. At its core, ThinkV2V builds on a practical MLLM-to-DiT architecture to turn explicit thinking over the source video and instruction into refined conditioning signals for video editing. Further, we equip it with a dedicated training and inference recipe, combining Progressive Curriculum Training, which gradually cultivates the model from basic editing to reasoning-intensive cases, with Inference-Time Thinking Scaling, which iteratively refines candidate prompts and selects the most reliable one, to better elicit reasoning in challenging editing scenarios. We also curate the ThinkV2V-150K dataset and introduce ThinkV2V-Bench to support training and evaluation of video editing with implicit intent and causal reasoning. Experimental results demonstrate the state-of-the-art performance of ThinkV2V on both complex and standard editing scenarios, in which our 5B-scale DiT model substantially outperforms larger 10B-scale baselines.