🤖 AI Summary
This study addresses the scarcity of large-scale paired data and the prohibitive annotation and synthesis costs in video editing by proposing a novel framework that transfers image editing capabilities to videos via context modeling. Methodologically, it introduces position encoding within a shared spatial coordinate system to enable pixel-level appearance propagation and test-time scaling. Furthermore, it constructs a unified diffusion objective, contextual visual demonstrations, and a chainable composable subtask architecture that decomposes complex edits into modular pipelines. Experimental results demonstrate that the proposed method achieves state-of-the-art performance on the OpenVE-Bench benchmark, thoroughly validating the effectiveness of each core component.
📝 Abstract
Building a capable video editor remains significantly harder than a video generator: editing requires (source, instruction, edited) triplets that are prohibitively expensive to annotate and difficult to synthesize at scale, whereas image editing has already reached maturity with millions of such pairs readily available. In this work, we introduce VINCIE-NExT, a unified framework that transfers editing capability from images to videos through in-context visual demonstrations, alleviating the need for large-scale paired video editing data. VINCIE-NExT decomposes video editing into a structured chain of composable sub-tasks (Video -> Image -> Image -> Video), routing editing intent through the image domain and enabling scalable joint training from heterogeneous image and video corpora under a unified diffusion objective. An image editing pair, synthesized by the model or supplied by the user, is prepended as an in-context visual demonstration that serves as a spatial appearance blueprint for every output frame. To ground appearance edits across the interleaved context, we introduce a novel position encoding that links image demonstrations and video frames in a shared spatial coordinate system, enabling pixel-faithful propagation of appearance changes to every output frame. Chain-of-Editing further provides principled test-time scaling: by executing the sub-task chain as progressive diffusion stages, editing quality can be improved by investing additional compute without retraining. Comprehensive experiments on OpenVE-Bench demonstrate the state-of-the-art performance across diverse editing categories, with ablations confirming the effectiveness of each component.