Score
Design and implement methods that synthesize intermediate video frames from sparse keyframes by estimating and exploiting motion cues to preserve temporal coherence, using motion interpolation and warping to guide pixel- or feature-level interpolation. Build edge-aware refinement and lightweight temporal reconstruction modules that maintain sharp boundaries and temporal stability while running efficiently on constrained/edge hardware.
Video Frame Interpolation (VFI) aims to synthesize high-fidelity intermediate frames, with core challenges including large motions, occlusions, illumination variations, and modeling of nonlinear motion. This work presents a systematic survey of over 250 publications and introduces, for the first time, a comprehensive methodological taxonomy for VFI—explicitly distinguishing Continuous-Time Frame Interpolation (CTFI) from Arbitrary-Time Frame Interpolation (ATFI). It unifies major technical paradigms, including optical flow estimation, kernel prediction, generative adversarial networks (GANs), Transformers, Mamba architectures, and diffusion models. Furthermore, it establishes the first end-to-end research map covering datasets, loss functions, evaluation metrics, and cross-domain applications. The resulting framework constitutes the most complete knowledge structure for VFI to date, providing standardized terminology, reproducible benchmarks, and a clear technological evolution roadmap—thereby significantly advancing the systematic development of low-level vision foundational tasks.
This work addresses the challenges of motion blur, structural distortion, and temporal inconsistency in video frame interpolation under large temporal gaps and complex motion. The authors propose a novel approach that leverages pre-trained image-to-video diffusion models without requiring retraining. By incorporating high-temporal-resolution motion cues from event cameras, they design a lightweight adapter architecture that fuses Image Warped Events (IWEs) with bidirectional sparse optical flow to generate spatiotemporally aligned structural and motion guidance signals, which are injected into the latent diffusion model. This method represents the first effective integration of event streams with pre-trained DiT-based video generation models, achieving state-of-the-art performance on both real-world and synthetic benchmarks, significantly improving interpolation fidelity and temporal coherence.
To address motion discontinuity and poor temporal consistency in keyframe-based video interpolation, this paper proposes a lightweight bidirectional diffusion sampling framework. Without retraining large-scale models, it fine-tunes pre-trained image-to-video diffusion models (e.g., Sora-like architectures) to enable bidirectional temporal modeling. The method initiates collaborative sampling from both end keyframes and introduces an overlapping estimation fusion strategy to enhance motion plausibility and structural fidelity of intermediate frames. To our knowledge, this is the first work to efficiently adapt unidirectional image-to-video diffusion models for keyframe interpolation. Extensive experiments demonstrate that our approach significantly outperforms optical-flow-based methods and existing diffusion-based interpolation techniques across multiple benchmarks, achieving state-of-the-art performance in visual quality, motion smoothness, and temporal consistency.
Existing generative video frame interpolation methods are constrained by fixed interpolation factors, limiting flexible control over output frame rate and temporal duration. To address this, we propose the first framework capable of synthesizing videos at arbitrary timestamps and arbitrary lengths. Our approach introduces timestamp-aware rotary positional encoding (TaRoPE) to enable precise temporal localization; designs a segmented conditional mechanism that decouples appearance and motion representations, ensuring long-term appearance consistency and motion coherence; and builds a multi-scale diffusion-based interpolation architecture. Evaluated on continuous interpolation tasks ranging from 2× to 32×, our method comprehensively outperforms state-of-the-art approaches, achieving significant improvements in visual quality and spatiotemporal continuity. Extensive experiments demonstrate strong generalization across complex real-world scenes, validating the robustness and flexibility of our framework.
Existing image-to-video generation methods suffer from limitations in camera trajectory control, temporal consistency, and geometric completeness. This work proposes an end-to-end framework based on dynamic 3D Gaussian splatting, which— for the first time—employs dynamic 3D Gaussian representations for single-image-driven video synthesis. The method jointly models camera motion and object dynamics within a single forward pass. By leveraging an explicit 3D scene representation, a motion sampling mechanism conditioned on a single input image, and differentiable rendering guided by prescribed camera trajectories, it achieves efficient, controllable, and temporally coherent video generation. Experiments on KITTI, Waymo, RealEstate10K, and DL3DV-10K demonstrate that the proposed approach significantly outperforms existing methods in both video quality and inference efficiency.
This work addresses the degradation of motion tweening quality caused by imprecise keyframe timing annotations. Existing approaches assume exact keyframe timestamps, yet practical annotations often contain temporal errors. We propose a robust motion sequence generation framework that jointly models an explicit time-warping function and spatial pose residuals. Our architecture incorporates a learnable temporal mapping module to dynamically retiming keyframes, while jointly optimizing temporal coherence and fine-grained pose details. Implemented as an end-to-end trainable generative neural network, the method requires no precise temporal priors. Extensive evaluation across multiple motion datasets demonstrates strong robustness to approximately timed keyframes: generated motions exhibit natural fluidity, plausible rhythm, and rich sub-motions—significantly outperforming fixed-timing baseline methods.
This work addresses the high memory consumption, inference latency, and detail degradation inherent in existing diffusion-based video frame interpolation methods that rely on multi-step sampling. To overcome these limitations, we propose SPEED—a single-step, pixel-level diffusion framework that jointly models multi-scale motion, structure, and appearance through a progressive multi-stage architecture with dynamic block scaling, directly predicting intermediate frames in pixel space. Key innovations include a Noise-Update-Only Attention mechanism that preserves semantic fidelity of conditioning frames while reducing computational overhead by nearly 50%, and a Drift-aware Timestep Sampling strategy that enhances single-step generation quality. Experiments demonstrate that SPEED achieves an 8.8% lower LPIPS on SNU-FILM, 63.3% faster inference, and 10.6% less memory usage; on 4K benchmarks, it improves LPIPS by up to 51.5% over prior methods.
Existing 4D Gaussian splatting methods suffer from overfitting to discrete time frames, hindering continuous-time reconstruction and leading to ghosting artifacts and temporal aliasing during interpolation. To address this, we propose RetimeGS, the first approach to explicitly model the temporal evolution of 3D Gaussians. Our method introduces optical flow–guided initialization and a triple-rendering supervision scheme that incorporates multi-view temporal consistency constraints, effectively mitigating temporal aliasing. RetimeGS achieves high-quality, ghosting-free, and temporally coherent continuous dynamic reconstruction, even under challenging conditions such as high-speed motion, non-rigid deformations, and severe occlusions, significantly outperforming current state-of-the-art techniques.
Autoregressive video generation remains impractical due to the high computational cost of iterative per-frame denoising, and existing cache reuse methods suffer from coarse granularity that fails to capture pixel-level motion dynamics. This work proposes MotionCache, a novel framework that introduces pixel-level motion awareness into the caching mechanism for the first time. Leveraging inter-frame differences as a lightweight motion proxy, MotionCache employs a coarse-to-fine strategy: it first establishes semantic consistency during a warm-up phase and then dynamically schedules cache update frequencies for individual tokens based on local motion intensity. Theoretical analysis reveals a critical link between cache-induced errors and residual instability. Experiments demonstrate significant acceleration—6.28× on SkyReels-V2 and 1.64× on MAGI-1—with negligible quality degradation of only 1% and 0.01% on VBench, respectively.
Traditional single-hypothesis optical flow often leads to ghosting, structural distortion, or blurriness in interpolated frames when confronted with ambiguous matching regions such as repetitive textures, symmetric structures, or motion blur. This work proposes the first multi-hypothesis optical flow estimation framework that operates without ground-truth flow supervision. By maintaining multiple candidate correspondences and employing a reliability-guided routing mechanism to select the optimal hypothesis, the method avoids soft blending and instead refines each hypothesis independently through anchor initialization and local attention. This approach significantly enhances interpolation accuracy, achieving state-of-the-art performance in terms of LPIPS and DISTS metrics on MA-HD and several standard video frame interpolation benchmarks, effectively mitigating ghosting artifacts and structural distortions.
This work addresses the challenge of preserving both fine detail fidelity and temporal motion consistency in video frame interpolation under high upscaling factors (e.g., 4×/8×) and high resolutions (2560×1440), where existing methods often suffer from structural distortions or motion incoherence. To this end, the authors propose FC-VFI, a novel approach leveraging a pre-trained video diffusion model to model temporal dependencies in latent space, thereby retaining structural details from the input frames. The method incorporates semantic matching lines to provide structure-aware motion guidance and introduces a temporal differential loss to enhance temporal consistency. Extensive experiments demonstrate that FC-VFI successfully interpolates 30 FPS videos to 120/240 FPS with superior visual quality, significantly outperforming state-of-the-art methods across diverse scenarios while maintaining high structural integrity and perceptual fidelity.