🤖 AI Summary
This study addresses the limitations of existing video editing methods, which struggle to simultaneously achieve semantic understanding and temporal consistency while incurring prohibitive computational costs. To overcome these challenges, this work proposes the AVE framework, which decouples video editing into two stages: keyframe editing and temporal propagation. Specifically, it leverages an anchoring mechanism to repurpose powerful image editors for processing keyframes, combined with a motion-guided diffusion model to enable training-free video generation. Extensive evaluations on benchmarks such as IVEBench demonstrate that the proposed approach achieves superior temporal consistency and content fidelity. Furthermore, it significantly reduces computational overhead while enhancing instruction-following capability. These results empirically validate a strong correlation between overall video editing quality and the performance of the underlying image editor.
📝 Abstract
Recent image editing systems have achieved impressive semantic understanding, visual fidelity, and instruction-following ability, while video editing remains substantially more difficult and costly. In this paper, we present a simple alternative to end-to-end video editing: instead of training a monolithic video editor, we transform a strong image editor into a video editor through anchor-based generation. Our key insight is that video editing can be decomposed into two subproblems: editing a sparse set of keyframes and propagating those edits across time. Based on this observation, we propose Anchor-based Video Editing (AVE), a two-stage framework in which a powerful image editor first performs composed editing on selected keyframes, and a motion-guided image-to-video diffusion model then generates the final video by treating the edited keyframes as fixed anchors. This design directly inherits the strengths of modern image editors while avoiding expensive end-to-end video editing training. Experiments on IVEBench and VIE-Bench show that AVE achieves strong performance in instruction following, temporal consistency, and content fidelity. Further ablations reveal that final video editing quality is strongly correlated with the quality of the image editor, suggesting that future progress in video editing may come from stronger image editing foundations and lightweight transfer to video. Code is available at https://github.com/wangf3014/AVE.