Transforming Image Editors into Video Editors

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing video editing methods, which struggle to simultaneously achieve semantic understanding and temporal consistency while incurring prohibitive computational costs. To overcome these challenges, this work proposes the AVE framework, which decouples video editing into two stages: keyframe editing and temporal propagation. Specifically, it leverages an anchoring mechanism to repurpose powerful image editors for processing keyframes, combined with a motion-guided diffusion model to enable training-free video generation. Extensive evaluations on benchmarks such as IVEBench demonstrate that the proposed approach achieves superior temporal consistency and content fidelity. Furthermore, it significantly reduces computational overhead while enhancing instruction-following capability. These results empirically validate a strong correlation between overall video editing quality and the performance of the underlying image editor.
📝 Abstract
Recent image editing systems have achieved impressive semantic understanding, visual fidelity, and instruction-following ability, while video editing remains substantially more difficult and costly. In this paper, we present a simple alternative to end-to-end video editing: instead of training a monolithic video editor, we transform a strong image editor into a video editor through anchor-based generation. Our key insight is that video editing can be decomposed into two subproblems: editing a sparse set of keyframes and propagating those edits across time. Based on this observation, we propose Anchor-based Video Editing (AVE), a two-stage framework in which a powerful image editor first performs composed editing on selected keyframes, and a motion-guided image-to-video diffusion model then generates the final video by treating the edited keyframes as fixed anchors. This design directly inherits the strengths of modern image editors while avoiding expensive end-to-end video editing training. Experiments on IVEBench and VIE-Bench show that AVE achieves strong performance in instruction following, temporal consistency, and content fidelity. Further ablations reveal that final video editing quality is strongly correlated with the quality of the image editor, suggesting that future progress in video editing may come from stronger image editing foundations and lightweight transfer to video. Code is available at https://github.com/wangf3014/AVE.
Problem

Research questions and friction points this paper is trying to address.

Video Editing
Image Editing
Temporal Consistency
Content Fidelity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Anchor-based Video Editing
Image-to-Video Diffusion
Keyframe Propagation
Video Editing Decomposition
Instruction Following
🔎 Similar Papers
No similar papers found.