video frame interpolation

Designs and implements algorithms or models that synthesize one or more intermediate image frames between given video frames, producing temporally consistent, high‑quality frames that respect motion, occlusion, and appearance changes. Work includes estimating motion or correspondence, handling occlusions and lighting/brightness variation, measuring interpolation accuracy and temporal smoothness, and integrating interpolation into frame‑rate conversion or slow‑motion pipelines.

videoframeinterpolation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.19
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Arbitrary Generative Video Interpolation

Oct 01, 2025
GZ
Guozhen Zhang
🏛️ Nanjing University | Tencent Hunyuan | Shanghai AI Laboratory

Existing generative video frame interpolation methods are constrained by fixed interpolation factors, limiting flexible control over output frame rate and temporal duration. To address this, we propose the first framework capable of synthesizing videos at arbitrary timestamps and arbitrary lengths. Our approach introduces timestamp-aware rotary positional encoding (TaRoPE) to enable precise temporal localization; designs a segmented conditional mechanism that decouples appearance and motion representations, ensuring long-term appearance consistency and motion coherence; and builds a multi-scale diffusion-based interpolation architecture. Evaluated on continuous interpolation tasks ranging from 2× to 32×, our method comprehensively outperforms state-of-the-art approaches, achieving significant improvements in visual quality and spatiotemporal continuity. Extensive experiments demonstrate strong generalization across complex real-world scenes, validating the robustness and flexibility of our framework.

Enables video interpolation at arbitrary timestamps and lengthsEnsures spatiotemporal continuity across variable-length generated segmentsOvercomes fixed-frame limitations in generative video interpolation methods

Efficient motion-based metrics for video frame interpolation

Aug 12, 2025
CD
Conall Daly
🏛️ Sigmedia Group | Trinity College Dublin

Evaluating the perceptual quality of intermediate frames generated by video frame interpolation remains challenging, particularly due to the limitations of pixel-level fidelity metrics in capturing visual comfort and motion smoothness. Method: This paper proposes a novel no-reference quality metric based on optical flow field divergence, explicitly modeling motion smoothness and perceptual comfort through a motion consistency measure—bypassing reliance on pixel-wise reconstruction error. The metric is calibrated via regression against subjective scores from the BVI-VFI dataset. Contribution/Results: Compared to FloLPIPS, the proposed metric achieves a 2.7× speedup in computation while attaining a PLCC of 0.51 with subjective ratings—significantly outperforming PSNR and SSIM. It reliably identifies interpolated frames exhibiting low distortion yet high visual comfort. To our knowledge, this is the first work to explicitly incorporate flow field divergence into frame interpolation quality assessment, striking a superior balance between perceptual consistency and computational efficiency, thereby providing a more reliable evaluation benchmark for advanced interpolation algorithms.

Assessing perceptual quality of interpolated video framesDeveloping efficient motion-based video quality metricsEvaluating frame interpolation algorithms using motion divergence

Magic Fixup: Streamlining Photo Editing by Watching Dynamic Videos

Mar 19, 2024
HA
Hadi Alzayer
🏛️ Adobe | University of Maryland, College Park

This work addresses the challenge of synthesizing high-fidelity, photorealistic images from coarse layout edits. To mitigate second-order artifacts—including illumination mismatch, missing shadows, and physically implausible object interactions—the authors propose a diffusion-based inpainting method leveraging video temporal modeling. The method introduces a novel dual-motion modeling mechanism—optical-flow-guided warping coupled with hierarchical feature injection—supervised by paired video frames, enabling joint optimization of layout alignment, illumination consistency, and physically grounded object interactions. By integrating a pre-trained diffusion model, layout-constrained fine-tuning, and a dynamically constructed video dataset, the approach achieves fine-grained detail transfer and multi-factor coherent generation. Experiments demonstrate significant improvements in output photorealism, geometric consistency, and scene plausibility, while preserving object identity and texture fidelity.

Generates photorealistic images from coarse editsTransfers fine details while adapting to new lightingUses video supervision for realistic object interactions

A new dataset and comparison for multi-camera frame synthesis

Aug 12, 2025
CD
Conall Daly
🏛️ Sigmedia Group | Trinity College Dublin

Existing frame interpolation and novel view synthesis methods suffer from distributional mismatches in training data—frame interpolation focuses on temporal motion from a single camera, while view synthesis targets stereo depth estimation—preventing fair cross-task comparison. To address this, we introduce the first dense linear camera array dataset explicitly designed for multi-view frame generation, enabling unified evaluation across both temporal and spatial dimensions and filling a critical gap in cross-modal video generation benchmarks. Leveraging this dataset, we conduct a systematic benchmark of 3D Gaussian Splatting, classical optical flow-based methods, and deep learning-based frame interpolation algorithms. Results reveal a performance reversal: on real-world scenes, traditional methods outperform deep learning approaches by ~3.5 dB PSNR; conversely, on synthetic scenes, 3D Gaussian Splatting surpasses others by nearly 5 dB. This work establishes a new empirical standard for evaluating video generation models.

Assess performance on real vs synthetic image dataCompare frame interpolation and view synthesis methodsDevelop multi-camera dataset for fair evaluation

Pathways on the Image Manifold: Image Editing via Video Generation

Nov 25, 2024
NR
Noam Rotstein
🏛️ Technion - Israel Institute of Technology

Existing image diffusion models exhibit low accuracy in complex text-guided editing and often degrade critical content of the original image. To address this, we reformulate static image editing as a temporal evolution process and, for the first time, leverage a pre-trained image-to-video diffusion model to synthesize a manifold-continuous transition path—from source to edited image—along an implicit temporal trajectory. This temporal consistency enforces spatial semantic coherence without requiring fine-tuning or additional training. Our approach integrates manifold-constrained optimization and implicit temporal modeling to jointly preserve editing fidelity and structural/identity integrity of the input. Evaluated on text-driven image editing, our method achieves state-of-the-art performance, outperforming leading approaches both quantitatively (e.g., higher CLIP-Score, lower LPIPS) and qualitatively (e.g., sharper details, better semantic alignment, and stronger identity preservation).

Improving accuracy in complex image editsPreserving key elements of original imagesUsing video models for consistent image editing

Latest Papers

What's happening recently
View more

This work explores how to achieve video frame interpolation using only an image foundation model with spatial editing capabilities, without introducing explicit temporal modeling or motion estimation modules. By applying parameter-efficient fine-tuning via Low-Rank Adaptation (LoRA) to the pre-trained Qwen-Image-Edit model, the method activates its latent temporal reasoning ability with merely 64–256 training samples. This study is the first to demonstrate that static image editing models inherently possess transferable temporal understanding, enabling cross-modal generalization from spatial editing to video interpolation. Notably, this approach achieves data-efficient video synthesis without any architectural modifications, offering a novel paradigm particularly suitable for resource-constrained scenarios.

Few-Shot LearningFoundation ModelsImage Editing

Beyond the Visible: Disocclusion-Aware Editing via Proxy Dynamic Graphs

Dec 15, 2025
AQ
Anran Qi
🏛️ Inria | Université Côte d’Azur | University of Edinburgh | Adobe Research | UCL

Existing image-to-video methods struggle to simultaneously achieve motion controllability and content editability in disoccluded regions. To address this, we propose the Proxy Dynamic Graph (PDG), the first framework to explicitly model the decoupled relationship between visibility and motion, enabling training-free, inference-time controllable generation. PDG employs a lightweight graph structure to drive part-wise motion, integrating a frozen diffusion prior with motion-flow-guided, visibility-aware latent synthesis—thereby unifying loose pose editing and precise appearance specification. Our method significantly outperforms state-of-the-art approaches on articulated scenes—including furniture, vehicles, and deformable objects—while enabling accurate appearance editing in disoccluded regions. Generated videos exhibit physically plausible motion structures, high visual consistency across frames, and strong user controllability without fine-tuning.

Enforcing user-specified content in newly revealed disoccluded areas.Generating controllable motion in image-to-video synthesis.Separating motion specification from appearance synthesis for predictability.

This work addresses the challenge of controllably editing the motion trajectory of a target object in videos while preserving the original scene content. To this end, the authors propose a two-stage framework: first, a cross-view motion transformation module maps a user-specified trajectory—provided only in the initial frame—into per-frame bounding boxes that account for camera motion; second, a motion-conditioned video resynthesis module generates the object along this trajectory while maintaining background consistency. By eliminating the need for complex point-trajectory inputs, the method significantly enhances user-friendliness and temporal coherence. Experiments demonstrate that the approach produces more realistic, temporally consistent, and controllable motion edits on diverse real-world videos compared to existing image-to-video or video-to-video methods.

camera motionmotion pathobject motion editing

This work addresses the challenges of poor motion controllability, low perceptual quality, and temporal inconsistency in video frame interpolation by proposing a training-free interpolation framework. It leverages a pretrained optical flow model to construct symmetric nonlinear motion-guided frames, which serve as latent-space priors to iteratively steer a pretrained video diffusion model for high-fidelity and motion-coherent intermediate frame synthesis. The method innovatively integrates symmetric nonlinear motion modeling with a pretrained video diffusion model and introduces a confidence map fusion mechanism that balances structural reliability and textural realism in ambiguous regions such as occlusions and object boundaries. Extensive experiments on standard benchmarks—including DAVIS, Sintel, and KITTI—demonstrate superior performance in perceptual quality, reconstruction accuracy, and temporal consistency.

Diffusion ModelsMotion CorrespondenceOptical Flow

Beyond Boundary Frames: Audio-Visual Semantic Guidance for Context-Aware Video Interpolation

Dec 03, 2025
YD
Yuchen Deng
🏛️ Shenzhen International Graduate School, Tsinghua University | Pengcheng Laboratory

Video frame interpolation suffers from inaccurate motion modeling and temporal inconsistency when handling fast, complex nonlinear motions—particularly critical in fine-grained tasks like audio-video synchronization. To address this, we propose a context-aware, multimodal-guidable interpolation framework. Methodologically, we adopt a DiT backbone and design a decoupled multimodal fusion mechanism supporting conditional inputs including text, audio, images, and videos. We introduce start-end frame difference embeddings to modulate sampling and loss computation, and employ a dynamically adjusted progressive multi-stage training strategy to enhance fine-grained motion modeling while preserving core generative capabilities. Experiments demonstrate that our method outperforms state-of-the-art approaches on both general frame interpolation and audio-video synchronized interpolation, achieving significant improvements in motion accuracy and temporal consistency. These results validate its effectiveness in cross-modal collaborative motion modeling and strong generalization across diverse modalities.

Handles fast, complex, non-linear motion in video interpolationImproves sharpness and consistency in audio-visual synchronized tasksUnifies multi-modal conditioning for diverse interpolation scenarios

Hot Scholars

MG

Minyi Guo

IEEE Fellow, Chair Professor, Shanghai Jiao Tong University
Parallel ComputingCompiler OptimizationCloud ComputingNetworking
HG

Harsh Goel

University of Texas at Austin
Reinforcement LearningRoboticsGenerative AINeurosymbolic AI
HP

HyunWook Park

Professor of Electrical Engineering, KAIST
Image processingMedical imaging
XZ

Xiawu Zheng

Associate Professor, IEEE Senior Member, Xiamen University
Automated Machine LearningNetwork CompressionNeural Architecture SearchAutoML