🤖 AI Summary
Existing video editing methods predominantly focus on appearance modification while lacking effective control over dynamics. This work proposes the RVD framework, which introduces a novel mechanism to decouple dynamic tokens from visual context. By employing self-supervised reconstruction to disentangle motion and appearance information, combined with a language-guided editor and a counterfactual video pair training pipeline, the method enables motion-only adjustments while preserving visual content. Built upon a Transformer architecture with a two-stage training strategy, RVD achieves training-free video speed manipulation, temporal reordering, and controllable appearance re-rendering. Consequently, this approach significantly enhances both the flexibility and efficiency of video dynamic editing.
📝 Abstract
Most video editing methods focus on changing the appearance of the source video, while offering limited control over its dynamics. We introduce Reimagine Video Dynamics (RVD), a framework that disentangles a compact, editable dynamics token from visual context. We learn this token through self-supervised reconstruction: given the first frame as visual context, a renderer must recover the original video from the dynamics token, encouraging it to capture how the scene evolves rather than how it looks. This disentanglement allows video dynamics to be edited directly while preserving visual context. We develop a language-guided dynamics-token editor that transforms source dynamics into target dynamics, and train it with a scalable counterfactual video-pair pipeline and a two-stage training strategy. Extensive experiments show that RVD enables effective video dynamics editing, training-free retiming, and appearance-controlled re-rendering.