🤖 AI Summary
Existing generative video frame interpolation methods are constrained by fixed interpolation factors, limiting flexible control over output frame rate and temporal duration. To address this, we propose the first framework capable of synthesizing videos at arbitrary timestamps and arbitrary lengths. Our approach introduces timestamp-aware rotary positional encoding (TaRoPE) to enable precise temporal localization; designs a segmented conditional mechanism that decouples appearance and motion representations, ensuring long-term appearance consistency and motion coherence; and builds a multi-scale diffusion-based interpolation architecture. Evaluated on continuous interpolation tasks ranging from 2× to 32×, our method comprehensively outperforms state-of-the-art approaches, achieving significant improvements in visual quality and spatiotemporal continuity. Extensive experiments demonstrate strong generalization across complex real-world scenes, validating the robustness and flexibility of our framework.
📝 Abstract
Video frame interpolation (VFI), which generates intermediate frames from given start and end frames, has become a fundamental function in video generation applications. However, existing generative VFI methods are constrained to synthesize a fixed number of intermediate frames, lacking the flexibility to adjust generated frame rates or total sequence duration. In this work, we present ArbInterp, a novel generative VFI framework that enables efficient interpolation at any timestamp and of any length. Specifically, to support interpolation at any timestamp, we propose the Timestamp-aware Rotary Position Embedding (TaRoPE), which modulates positions in temporal RoPE to align generated frames with target normalized timestamps. This design enables fine-grained control over frame timestamps, addressing the inflexibility of fixed-position paradigms in prior work. For any-length interpolation, we decompose long-sequence generation into segment-wise frame synthesis. We further design a novel appearance-motion decoupled conditioning strategy: it leverages prior segment endpoints to enforce appearance consistency and temporal semantics to maintain motion coherence, ensuring seamless spatiotemporal transitions across segments. Experimentally, we build comprehensive benchmarks for multi-scale frame interpolation (2x to 32x) to assess generalizability across arbitrary interpolation factors. Results show that ArbInterp outperforms prior methods across all scenarios with higher fidelity and more seamless spatiotemporal continuity. Project website: https://mcg-nju.github.io/ArbInterp-Web/.