temporal feature propagation

Designs and evaluates methods that propagate feature representations across neighboring video frames using information from recent past frames to capture short-term motion cues and enforce temporal consistency. Focuses on lightweight, low-compute propagation mechanisms (e.g., warping or recurrent updates and other flow-free approaches) that avoid costly components such as dense temporal attention or explicit optical-flow computation.

temporalfeaturepropagation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.15
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Training-Free Motion-Guided Video Generation with Enhanced Temporal Consistency Using Motion Consistency Loss

Jan 13, 2025
XZ
Xinyu Zhang
🏛️ The University of Adelaide | The University of New South Wales

This work addresses the problem of training-free motion-guided video generation. We propose a zero-shot motion-consistency optimization method based on diffusion models. Our core contribution is the first motion-consistency loss function, which explicitly models inter-frame correlations of a reference video within intermediate feature layers of the diffusion model. By backpropagating gradients through this loss in latent space, the method guides the initial noise sampling process to implicitly learn and reproduce the target motion pattern—without fine-tuning or additional training. The entire optimization requires only a single forward–backward pass. Experiments demonstrate significant improvements in temporal coherence and motion fidelity across diverse motion control tasks, establishing a new state-of-the-art benchmark for training-free video generation.

Natural Motion AdaptationTraining EfficiencyVideo Generation

Temporal Realism Evaluation of Generated Videos Using Compressed-Domain Motion Vectors

Nov 17, 2025
MO
Mert Onur Cakiroglu
🏛️ Indiana University Bloomington | Hamad Bin Khalifa University

Current generative video models suffer from insufficient temporal realism, while mainstream evaluation metrics exhibit low sensitivity to motion modeling. To address this, we propose the first temporal fidelity assessment framework based on compressed-domain motion vectors (MVs) extracted from H.264/HEVC bitstreams. Leveraging MV statistical properties—including motion entropy and optical flow field structure—we quantify dynamic behavioral discrepancies between generated and real videos. We innovatively employ KL divergence, Jensen–Shannon divergence, and Wasserstein distance to measure MV distributional differences, and design three MV-RGB fusion mechanisms—channel concatenation, cross-attention, and joint embedding—to enhance temporal modeling. Evaluated on GenVidBench across eight state-of-the-art generators, our method enables fine-grained assessment (with Pika and SVD achieving top performance). When integrated with MV features, ResNet and I3D achieve 99.0% accuracy in binary fake/real video classification.

Evaluating temporal realism in generated videos using motion vectorsIdentifying motion defects in generative models via compressed-domain analysisImproving video classification accuracy with motion-aware fusion techniques

This work addresses the high computational cost, substantial memory consumption, and suboptimal generation quality in video continuation tasks by proposing an efficient flow-matching-based generative approach. The method fine-tunes a pretrained text-to-video diffusion model to learn a vector field mapping current frames to subsequent ones, directly modeling inter-frame flow trajectories without introducing noise and with reduced input dimensionality. It innovatively incorporates intrinsic optimal coupling and target inversion mechanisms to straighten flow paths and enhance frame-wise correspondence accuracy. Remarkably, with only five neural function evaluations, the proposed method achieves significant improvements over existing approaches in both FID and FVD metrics, simultaneously boosting generation efficiency and visual fidelity.

fast generationflow modelmemory efficiency

Feed-Forward Bullet-Time Reconstruction of Dynamic Scenes from Monocular Videos

Dec 04, 2024
HL
Hanxue Liang
🏛️ NVIDIA | University of Cambridge | MIT | Nanyang Technological University | University of Toronto | Vector Institute

Existing static feed-forward scene reconstruction methods suffer from poor generalization and fail to model dynamic content effectively. To address this, we propose the first motion-aware feed-forward framework for dynamic scene reconstruction, enabling real-time bullet-time rendering and novel-view synthesis from monocular video input. Our approach employs a 3D Gaussian splatting representation integrated with a cross-frame spatiotemporal aggregation mechanism, jointly modeling static backgrounds and dynamic foregrounds without iterative optimization. The model processes monocular video end-to-end and reconstructs the entire scene within 150 ms—significantly outperforming optimization-based methods in speed. It achieves state-of-the-art performance on both static and dynamic benchmarks, delivering strong generalization, high-fidelity reconstruction, and millisecond-level inference latency.

Enables high-quality novel view synthesis using 3D Gaussian SplattingImproves generalization across diverse static and dynamic environmentsReconstructs dynamic scenes from monocular videos in real-time

Looking Backward: Streaming Video-to-Video Translation with Feature Banks

May 24, 2024
FL
Feng Liang
🏛️ UT Austin | UC Berkeley

To address the challenge of simultaneously achieving infinite-frame processing and temporal consistency in real-time streaming video-to-video (V2V) translation, this paper introduces the first diffusion-based V2V architecture designed explicitly for streaming scenarios. Methodologically, we propose a backward-looking feature bank mechanism that dynamically stores historical features and directly fuses them into self-attention computation, thereby extending cross-frame attention without requiring model fine-tuning—enabling plug-and-play integration with existing image diffusion models. Our contributions are fourfold: (1) establishing the first streaming-aware V2V diffusion paradigm; (2) introducing a novel feature-bank-based temporal modeling mechanism; (3) achieving 20 FPS on a single A100 GPU—15× to 158× faster than FlowVid; and (4) demonstrating significant improvements in temporal consistency through both quantitative evaluation and user studies.

Real-time video-to-video translationStreaming frame processingTemporal consistency maintenance

Latest Papers

What's happening recently
View more

This work addresses the lack of a universal, model-agnostic geometric constraint in existing optical flow learning methods, which often leads to inconsistent performance across varying supervision schemes and data configurations. The authors propose trilinear consistency as a fundamental geometric prior: given any two optical flow fields, the third can be derived through composition, and consistency among all three is explicitly enforced. This constraint is applicable across diverse scenarios—including image pairs, multi-frame videos, and synthetically transformed sequences—and is introduced for the first time in a plug-and-play manner that is independent of network architecture, supervision type, or additional annotations. It seamlessly integrates with existing methods for joint optimization. Extensive experiments demonstrate consistent performance gains under supervised, unsupervised, and transfer learning settings, with negligible computational overhead.

flow consistencygeometric constraintoptical flow

This work addresses geometric inconsistencies in text-to-video generation—such as object deformation, texture drift, and non-rigid background motion—by introducing a geometric consistency reward mechanism that explicitly optimizes temporal geometric structure during reinforcement fine-tuning of diffusion models. For the first time, geometric consistency is formulated as a directly optimizable objective without modifying the model’s latent space, making the approach applicable to complex dynamic scenes involving both camera and object motion. By integrating optical flow, depth-pose estimation, and feature correspondence techniques, the method effectively disentangles rigid background from dynamic object regions and evaluates their consistency separately. Experiments demonstrate that this approach substantially reduces temporal geometric artifacts while preserving high visual fidelity, outperforming strong existing baselines.

camera motiongeometric consistencyobject deformation

Existing video frame interpolation methods often suffer from motion drift, directional ambiguity, and boundary misalignment due to unidirectional generation, and they lack temporal consistency over long sequences. This work proposes a bidirectionally cycle-consistent video diffusion interpolation framework that employs learnable directional tokens to guide a shared backbone network, jointly optimizing forward synthesis and backward reconstruction within a unified architecture to achieve logically invertible motion trajectories. During training, bidirectional cycle consistency is enforced as a regularizer, complemented by a curriculum learning strategy that progressively optimizes from short to long sequences. At inference, the model requires only a single forward pass. The proposed method significantly outperforms strong baselines on 37- and 73-frame interpolation tasks, achieving state-of-the-art performance in image quality, motion smoothness, and dynamic control without incurring additional computational overhead.

boundary misalignmentdirectional ambiguitymotion drift

Traditional single-hypothesis optical flow often leads to ghosting, structural distortion, or blurriness in interpolated frames when confronted with ambiguous matching regions such as repetitive textures, symmetric structures, or motion blur. This work proposes the first multi-hypothesis optical flow estimation framework that operates without ground-truth flow supervision. By maintaining multiple candidate correspondences and employing a reliability-guided routing mechanism to select the optimal hypothesis, the method avoids soft blending and instead refines each hypothesis independently through anchor initialization and local attention. This approach significantly enhances interpolation accuracy, achieving state-of-the-art performance in terms of LPIPS and DISTS metrics on MA-HD and several standard video frame interpolation benchmarks, effectively mitigating ghosting artifacts and structural distortions.

correspondence estimationmatching ambiguitymultiple hypothesis

Hot Scholars

HL

Haotong Lin

Zhejiang university
Computer Vision and Graphics
SP

Sida Peng

Zhejiang University
Computer VisionComputer Graphics
HL

Huaqiu Li

Tsinghua University
computer visionmachine learning
MK

Matej Kristan

Full Professor at Faculty of computer and information science, University of Ljubljana
Computer visionMachine learningPattern recognition
CW

Cong Wan

Xian Jiaotong University
AIGC3Ddiffusion