Score
Designs and evaluates methods that propagate feature representations across neighboring video frames using information from recent past frames to capture short-term motion cues and enforce temporal consistency. Focuses on lightweight, low-compute propagation mechanisms (e.g., warping or recurrent updates and other flow-free approaches) that avoid costly components such as dense temporal attention or explicit optical-flow computation.
This work addresses the problem of training-free motion-guided video generation. We propose a zero-shot motion-consistency optimization method based on diffusion models. Our core contribution is the first motion-consistency loss function, which explicitly models inter-frame correlations of a reference video within intermediate feature layers of the diffusion model. By backpropagating gradients through this loss in latent space, the method guides the initial noise sampling process to implicitly learn and reproduce the target motion pattern—without fine-tuning or additional training. The entire optimization requires only a single forward–backward pass. Experiments demonstrate significant improvements in temporal coherence and motion fidelity across diverse motion control tasks, establishing a new state-of-the-art benchmark for training-free video generation.
Current generative video models suffer from insufficient temporal realism, while mainstream evaluation metrics exhibit low sensitivity to motion modeling. To address this, we propose the first temporal fidelity assessment framework based on compressed-domain motion vectors (MVs) extracted from H.264/HEVC bitstreams. Leveraging MV statistical properties—including motion entropy and optical flow field structure—we quantify dynamic behavioral discrepancies between generated and real videos. We innovatively employ KL divergence, Jensen–Shannon divergence, and Wasserstein distance to measure MV distributional differences, and design three MV-RGB fusion mechanisms—channel concatenation, cross-attention, and joint embedding—to enhance temporal modeling. Evaluated on GenVidBench across eight state-of-the-art generators, our method enables fine-grained assessment (with Pika and SVD achieving top performance). When integrated with MV features, ResNet and I3D achieve 99.0% accuracy in binary fake/real video classification.
This work addresses the high computational cost, substantial memory consumption, and suboptimal generation quality in video continuation tasks by proposing an efficient flow-matching-based generative approach. The method fine-tunes a pretrained text-to-video diffusion model to learn a vector field mapping current frames to subsequent ones, directly modeling inter-frame flow trajectories without introducing noise and with reduced input dimensionality. It innovatively incorporates intrinsic optimal coupling and target inversion mechanisms to straighten flow paths and enhance frame-wise correspondence accuracy. Remarkably, with only five neural function evaluations, the proposed method achieves significant improvements over existing approaches in both FID and FVD metrics, simultaneously boosting generation efficiency and visual fidelity.
Existing static feed-forward scene reconstruction methods suffer from poor generalization and fail to model dynamic content effectively. To address this, we propose the first motion-aware feed-forward framework for dynamic scene reconstruction, enabling real-time bullet-time rendering and novel-view synthesis from monocular video input. Our approach employs a 3D Gaussian splatting representation integrated with a cross-frame spatiotemporal aggregation mechanism, jointly modeling static backgrounds and dynamic foregrounds without iterative optimization. The model processes monocular video end-to-end and reconstructs the entire scene within 150 ms—significantly outperforming optimization-based methods in speed. It achieves state-of-the-art performance on both static and dynamic benchmarks, delivering strong generalization, high-fidelity reconstruction, and millisecond-level inference latency.
To address the challenge of simultaneously achieving infinite-frame processing and temporal consistency in real-time streaming video-to-video (V2V) translation, this paper introduces the first diffusion-based V2V architecture designed explicitly for streaming scenarios. Methodologically, we propose a backward-looking feature bank mechanism that dynamically stores historical features and directly fuses them into self-attention computation, thereby extending cross-frame attention without requiring model fine-tuning—enabling plug-and-play integration with existing image diffusion models. Our contributions are fourfold: (1) establishing the first streaming-aware V2V diffusion paradigm; (2) introducing a novel feature-bank-based temporal modeling mechanism; (3) achieving 20 FPS on a single A100 GPU—15× to 158× faster than FlowVid; and (4) demonstrating significant improvements in temporal consistency through both quantitative evaluation and user studies.
This work addresses the lack of a universal, model-agnostic geometric constraint in existing optical flow learning methods, which often leads to inconsistent performance across varying supervision schemes and data configurations. The authors propose trilinear consistency as a fundamental geometric prior: given any two optical flow fields, the third can be derived through composition, and consistency among all three is explicitly enforced. This constraint is applicable across diverse scenarios—including image pairs, multi-frame videos, and synthetically transformed sequences—and is introduced for the first time in a plug-and-play manner that is independent of network architecture, supervision type, or additional annotations. It seamlessly integrates with existing methods for joint optimization. Extensive experiments demonstrate consistent performance gains under supervised, unsupervised, and transfer learning settings, with negligible computational overhead.
This work addresses geometric inconsistencies in text-to-video generation—such as object deformation, texture drift, and non-rigid background motion—by introducing a geometric consistency reward mechanism that explicitly optimizes temporal geometric structure during reinforcement fine-tuning of diffusion models. For the first time, geometric consistency is formulated as a directly optimizable objective without modifying the model’s latent space, making the approach applicable to complex dynamic scenes involving both camera and object motion. By integrating optical flow, depth-pose estimation, and feature correspondence techniques, the method effectively disentangles rigid background from dynamic object regions and evaluates their consistency separately. Experiments demonstrate that this approach substantially reduces temporal geometric artifacts while preserving high visual fidelity, outperforming strong existing baselines.
Existing video frame interpolation methods often suffer from motion drift, directional ambiguity, and boundary misalignment due to unidirectional generation, and they lack temporal consistency over long sequences. This work proposes a bidirectionally cycle-consistent video diffusion interpolation framework that employs learnable directional tokens to guide a shared backbone network, jointly optimizing forward synthesis and backward reconstruction within a unified architecture to achieve logically invertible motion trajectories. During training, bidirectional cycle consistency is enforced as a regularizer, complemented by a curriculum learning strategy that progressively optimizes from short to long sequences. At inference, the model requires only a single forward pass. The proposed method significantly outperforms strong baselines on 37- and 73-frame interpolation tasks, achieving state-of-the-art performance in image quality, motion smoothness, and dynamic control without incurring additional computational overhead.
研究通过分析V-JEPA 2和VideoMAE-v2模型,探讨了视频基础模型中时空表示的编码内容、出现位置及几何组织方式,并使用轻量级探针来发现三种时间属性。
Traditional single-hypothesis optical flow often leads to ghosting, structural distortion, or blurriness in interpolated frames when confronted with ambiguous matching regions such as repetitive textures, symmetric structures, or motion blur. This work proposes the first multi-hypothesis optical flow estimation framework that operates without ground-truth flow supervision. By maintaining multiple candidate correspondences and employing a reliability-guided routing mechanism to select the optimal hypothesis, the method avoids soft blending and instead refines each hypothesis independently through anchor initialization and local attention. This approach significantly enhances interpolation accuracy, achieving state-of-the-art performance in terms of LPIPS and DISTS metrics on MA-HD and several standard video frame interpolation benchmarks, effectively mitigating ghosting artifacts and structural distortions.