Score
Designs and implements algorithms or models that synthesize one or more intermediate image frames between given video frames, producing temporally consistent, high‑quality frames that respect motion, occlusion, and appearance changes. Work includes estimating motion or correspondence, handling occlusions and lighting/brightness variation, measuring interpolation accuracy and temporal smoothness, and integrating interpolation into frame‑rate conversion or slow‑motion pipelines.
Video Frame Interpolation (VFI) aims to synthesize high-fidelity intermediate frames, with core challenges including large motions, occlusions, illumination variations, and modeling of nonlinear motion. This work presents a systematic survey of over 250 publications and introduces, for the first time, a comprehensive methodological taxonomy for VFI—explicitly distinguishing Continuous-Time Frame Interpolation (CTFI) from Arbitrary-Time Frame Interpolation (ATFI). It unifies major technical paradigms, including optical flow estimation, kernel prediction, generative adversarial networks (GANs), Transformers, Mamba architectures, and diffusion models. Furthermore, it establishes the first end-to-end research map covering datasets, loss functions, evaluation metrics, and cross-domain applications. The resulting framework constitutes the most complete knowledge structure for VFI to date, providing standardized terminology, reproducible benchmarks, and a clear technological evolution roadmap—thereby significantly advancing the systematic development of low-level vision foundational tasks.
Existing generative video frame interpolation methods are constrained by fixed interpolation factors, limiting flexible control over output frame rate and temporal duration. To address this, we propose the first framework capable of synthesizing videos at arbitrary timestamps and arbitrary lengths. Our approach introduces timestamp-aware rotary positional encoding (TaRoPE) to enable precise temporal localization; designs a segmented conditional mechanism that decouples appearance and motion representations, ensuring long-term appearance consistency and motion coherence; and builds a multi-scale diffusion-based interpolation architecture. Evaluated on continuous interpolation tasks ranging from 2× to 32×, our method comprehensively outperforms state-of-the-art approaches, achieving significant improvements in visual quality and spatiotemporal continuity. Extensive experiments demonstrate strong generalization across complex real-world scenes, validating the robustness and flexibility of our framework.
Evaluating the perceptual quality of intermediate frames generated by video frame interpolation remains challenging, particularly due to the limitations of pixel-level fidelity metrics in capturing visual comfort and motion smoothness. Method: This paper proposes a novel no-reference quality metric based on optical flow field divergence, explicitly modeling motion smoothness and perceptual comfort through a motion consistency measure—bypassing reliance on pixel-wise reconstruction error. The metric is calibrated via regression against subjective scores from the BVI-VFI dataset. Contribution/Results: Compared to FloLPIPS, the proposed metric achieves a 2.7× speedup in computation while attaining a PLCC of 0.51 with subjective ratings—significantly outperforming PSNR and SSIM. It reliably identifies interpolated frames exhibiting low distortion yet high visual comfort. To our knowledge, this is the first work to explicitly incorporate flow field divergence into frame interpolation quality assessment, striking a superior balance between perceptual consistency and computational efficiency, thereby providing a more reliable evaluation benchmark for advanced interpolation algorithms.
This work addresses the challenge of synthesizing high-fidelity, photorealistic images from coarse layout edits. To mitigate second-order artifacts—including illumination mismatch, missing shadows, and physically implausible object interactions—the authors propose a diffusion-based inpainting method leveraging video temporal modeling. The method introduces a novel dual-motion modeling mechanism—optical-flow-guided warping coupled with hierarchical feature injection—supervised by paired video frames, enabling joint optimization of layout alignment, illumination consistency, and physically grounded object interactions. By integrating a pre-trained diffusion model, layout-constrained fine-tuning, and a dynamically constructed video dataset, the approach achieves fine-grained detail transfer and multi-factor coherent generation. Experiments demonstrate significant improvements in output photorealism, geometric consistency, and scene plausibility, while preserving object identity and texture fidelity.
Existing frame interpolation and novel view synthesis methods suffer from distributional mismatches in training data—frame interpolation focuses on temporal motion from a single camera, while view synthesis targets stereo depth estimation—preventing fair cross-task comparison. To address this, we introduce the first dense linear camera array dataset explicitly designed for multi-view frame generation, enabling unified evaluation across both temporal and spatial dimensions and filling a critical gap in cross-modal video generation benchmarks. Leveraging this dataset, we conduct a systematic benchmark of 3D Gaussian Splatting, classical optical flow-based methods, and deep learning-based frame interpolation algorithms. Results reveal a performance reversal: on real-world scenes, traditional methods outperform deep learning approaches by ~3.5 dB PSNR; conversely, on synthetic scenes, 3D Gaussian Splatting surpasses others by nearly 5 dB. This work establishes a new empirical standard for evaluating video generation models.
Existing image diffusion models exhibit low accuracy in complex text-guided editing and often degrade critical content of the original image. To address this, we reformulate static image editing as a temporal evolution process and, for the first time, leverage a pre-trained image-to-video diffusion model to synthesize a manifold-continuous transition path—from source to edited image—along an implicit temporal trajectory. This temporal consistency enforces spatial semantic coherence without requiring fine-tuning or additional training. Our approach integrates manifold-constrained optimization and implicit temporal modeling to jointly preserve editing fidelity and structural/identity integrity of the input. Evaluated on text-driven image editing, our method achieves state-of-the-art performance, outperforming leading approaches both quantitatively (e.g., higher CLIP-Score, lower LPIPS) and qualitatively (e.g., sharper details, better semantic alignment, and stronger identity preservation).
This work explores how to achieve video frame interpolation using only an image foundation model with spatial editing capabilities, without introducing explicit temporal modeling or motion estimation modules. By applying parameter-efficient fine-tuning via Low-Rank Adaptation (LoRA) to the pre-trained Qwen-Image-Edit model, the method activates its latent temporal reasoning ability with merely 64–256 training samples. This study is the first to demonstrate that static image editing models inherently possess transferable temporal understanding, enabling cross-modal generalization from spatial editing to video interpolation. Notably, this approach achieves data-efficient video synthesis without any architectural modifications, offering a novel paradigm particularly suitable for resource-constrained scenarios.
Existing image-to-video methods struggle to simultaneously achieve motion controllability and content editability in disoccluded regions. To address this, we propose the Proxy Dynamic Graph (PDG), the first framework to explicitly model the decoupled relationship between visibility and motion, enabling training-free, inference-time controllable generation. PDG employs a lightweight graph structure to drive part-wise motion, integrating a frozen diffusion prior with motion-flow-guided, visibility-aware latent synthesis—thereby unifying loose pose editing and precise appearance specification. Our method significantly outperforms state-of-the-art approaches on articulated scenes—including furniture, vehicles, and deformable objects—while enabling accurate appearance editing in disoccluded regions. Generated videos exhibit physically plausible motion structures, high visual consistency across frames, and strong user controllability without fine-tuning.
This work addresses the challenge of controllably editing the motion trajectory of a target object in videos while preserving the original scene content. To this end, the authors propose a two-stage framework: first, a cross-view motion transformation module maps a user-specified trajectory—provided only in the initial frame—into per-frame bounding boxes that account for camera motion; second, a motion-conditioned video resynthesis module generates the object along this trajectory while maintaining background consistency. By eliminating the need for complex point-trajectory inputs, the method significantly enhances user-friendliness and temporal coherence. Experiments demonstrate that the approach produces more realistic, temporally consistent, and controllable motion edits on diverse real-world videos compared to existing image-to-video or video-to-video methods.
This work addresses the challenges of poor motion controllability, low perceptual quality, and temporal inconsistency in video frame interpolation by proposing a training-free interpolation framework. It leverages a pretrained optical flow model to construct symmetric nonlinear motion-guided frames, which serve as latent-space priors to iteratively steer a pretrained video diffusion model for high-fidelity and motion-coherent intermediate frame synthesis. The method innovatively integrates symmetric nonlinear motion modeling with a pretrained video diffusion model and introduces a confidence map fusion mechanism that balances structural reliability and textural realism in ambiguous regions such as occlusions and object boundaries. Extensive experiments on standard benchmarks—including DAVIS, Sintel, and KITTI—demonstrate superior performance in perceptual quality, reconstruction accuracy, and temporal consistency.
Video frame interpolation suffers from inaccurate motion modeling and temporal inconsistency when handling fast, complex nonlinear motions—particularly critical in fine-grained tasks like audio-video synchronization. To address this, we propose a context-aware, multimodal-guidable interpolation framework. Methodologically, we adopt a DiT backbone and design a decoupled multimodal fusion mechanism supporting conditional inputs including text, audio, images, and videos. We introduce start-end frame difference embeddings to modulate sampling and loss computation, and employ a dynamically adjusted progressive multi-stage training strategy to enhance fine-grained motion modeling while preserving core generative capabilities. Experiments demonstrate that our method outperforms state-of-the-art approaches on both general frame interpolation and audio-video synchronized interpolation, achieving significant improvements in motion accuracy and temporal consistency. These results validate its effectiveness in cross-modal collaborative motion modeling and strong generalization across diverse modalities.