optical flow estimation

Estimating dense pixelwise motion between frames, including methods to generate temporally consistent pseudo-labels and preserve stable background regions over long horizons. Also involves metrics and procedures to quantify physical inconsistency in videos without ground-truth references.

opticalflowestimation

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Existing video generation models often suffer from spatial distortions due to inconsistent 3D geometry, yet prevailing evaluation metrics such as FVD struggle to distinguish plausible motion dynamics from geometric inaccuracies. To address this gap, this work proposes the Spatial Geometric Consistency (SGC) metric, which segments static and dynamic regions in generated videos and enforces 3D consistency by partitioning the static background into spatially coherent subregions. For each subregion, SGC estimates depth and local camera poses, then quantifies geometric fidelity through multi-view pose divergence. SGC is the first metric capable of accurately detecting geometric distortions while preserving sensitivity to realistic dynamic content. It demonstrates superior robustness and discriminative power over existing methods on both real-world and synthetic videos, effectively addressing a critical limitation in current video generation evaluation frameworks.

3D spatial geometric consistencycamera pose divergencedynamic generated videos

GeCo: A Differentiable Geometric Consistency Metric for Video Generation

Dec 24, 2025
LG
Leslie Gu
🏛️ Harvard University | Google DeepMind | Massachusetts Institute of Technology

To address geometric distortions and occlusion inconsistencies—two prevalent artifacts in video generation arising from static scene modeling—this paper introduces GeCo, a differentiable geometric consistency metric. GeCo jointly detects both artifacts by integrating residual optical flow with monocular depth priors and performing differentiable geometric reprojection to generate dense, interpretable consistency maps. Its key innovation lies in enabling training-free, model-agnostic guided optimization without fine-tuning or architecture-specific adaptation—overcoming limitations of prior methods that require model-specific tuning or parameter updates. Extensive evaluation across multiple state-of-the-art video generation models demonstrates that GeCo significantly suppresses geometric distortions and improves both geometric fidelity and spatiotemporal coherence of generated videos. By providing a general-purpose, plug-and-play metric for post-hoc analysis and optimization, GeCo establishes a new paradigm for universal video generation refinement.

Benchmarks video models and reduces deformation artifactsDetects geometric deformation and occlusion artifactsProduces interpretable dense consistency maps

This work addresses geometric inconsistencies in text-to-video generation—such as object deformation, texture drift, and non-rigid background motion—by introducing a geometric consistency reward mechanism that explicitly optimizes temporal geometric structure during reinforcement fine-tuning of diffusion models. For the first time, geometric consistency is formulated as a directly optimizable objective without modifying the model’s latent space, making the approach applicable to complex dynamic scenes involving both camera and object motion. By integrating optical flow, depth-pose estimation, and feature correspondence techniques, the method effectively disentangles rigid background from dynamic object regions and evaluates their consistency separately. Experiments demonstrate that this approach substantially reduces temporal geometric artifacts while preserving high visual fidelity, outperforming strong existing baselines.

camera motiongeometric consistencyobject deformation

This work addresses the challenges of temporally inconsistent predictions—such as flickering—in human-centric dense video tasks under motion, occlusion, and illumination changes, compounded by the scarcity of multi-task paired video supervision. To this end, we propose a scalable, photorealistic synthetic human video generation method that, for the first time, provides both frame-level and sequence-level pixel-wise annotations, including depth, surface normals, and masks. Leveraging this data, we develop a unified Vision Transformer (ViT)-based dense prediction architecture that integrates CSE human geometric priors with a lightweight channel reweighting module. Our approach employs a two-stage training strategy—static pretraining followed by dynamic sequence fine-tuning—to jointly optimize spatial and temporal consistency. The method achieves state-of-the-art performance on THuman2.1 and Hi4D benchmarks and demonstrates strong generalization to in-the-wild real-world videos.

flickeringhuman-centric dense predictionpaired supervision

Long video generation suffers from distribution shift induced by frame expansion when adapting short-video-trained diffusion models, resulting in visual inconsistency and motion distortion. To address this, we propose a training-free decoupling framework: first, we identify that PCA precisely separates global appearance consistency from local motion intensity in videos; second, we design a progressive feature fusion strategy and an initial noise mean reuse mechanism to break the strong appearance-motion coupling. Our method is plug-and-play for mainstream diffusion models, requiring only statistical feature reuse and cosine-similarity guidance—no fine-tuning or parameter updates. Experiments demonstrate significant improvements in long-video visual quality and inter-frame appearance consistency, achieving state-of-the-art performance across multiple benchmarks.

Addressing distribution shifts in long video generation from short videosDecoupling appearance and motion using Principal Component AnalysisIntegrating local and global information for visual consistency

Latest Papers

What's happening recently
View more

Current generative video models suffer from “temporal hallucinations”—such as ambiguous motion speed and temporal instability—due to inconsistent real-world frame rates in training data. This work proposes Visual Chronometer, the first method to systematically define and quantify the physical frame rate (PhyFPS) alignment problem in video generation. By directly predicting the underlying physical timescale from visual dynamics without relying on unreliable metadata, our approach enables accurate temporal calibration. We introduce two benchmarks, PhyFPS-Bench-Real and PhyFPS-Bench-Gen, a controlled temporal resampling training strategy, and a deep learning–based PhyFPS prediction model. Experiments reveal that mainstream generative models exhibit significant PhyFPS misalignment, and that correcting this discrepancy substantially improves the perceptual naturalness and temporal consistency of generated videos.

chronometric hallucinationmotion speedphysical frame rate

This work addresses the temporal instability of SAM2 in video segmentation under weak prompts—such as sparse points in a single frame—which manifests as boundary flickering, object disappearance, and region fluctuations, thereby compromising downstream reliability. The authors propose a training-free, inference-time temporal probability smoothing method that aligns segmentation probability maps across adjacent frames using optical flow. By adaptively fusing these aligned maps based on forward-backward optical flow consistency and pixel-wise uncertainty derived from segmentation entropy, the approach significantly enhances temporal coherence—measured by motion-compensated IoU, boundary stability, object persistence, and area smoothness—without sacrificing spatial accuracy. Notably, it requires no model modification or retraining, is compatible with any SAM2-based interactive video segmentation system, and supports real-time deployment.

flickering boundariesobject dropouttemporal instability

Existing video diffusion models often suffer from object deformations or spatial drift due to the absence of explicit 3D structural constraints. This work proposes a self-supervised framework that, for the first time, incorporates geometric priors into video generation training in the form of preference pairs. By leveraging a foundation model for geometry estimation, the method generates dense 3D consistency signals as preference labels and employs Direct Preference Optimization (DPO) to guide the diffusion model toward learning more physically plausible and temporally coherent spatiotemporal distributions. Notably, the approach requires no manual annotations and significantly enhances temporal stability, physical realism, and motion coherence in generated videos, outperforming current state-of-the-art methods across multiple evaluation metrics.

3D consistencygeometric coherencespatial drift

Existing video generation models lack objective, quantitative evaluation of geometric consistency, making it difficult to diagnose their shortcomings in 3D structure and physically plausible motion. To address this gap, this work proposes PDI-Bench, the first quantifiable benchmark framework tailored for geometric consistency assessment. It leverages SAM 2, MegaSaM, and CoTracker3 for object segmentation and point tracking, integrating monocular 3D reconstruction with projective geometry residual analysis to evaluate scale-depth alignment, 3D motion consistency, and structural rigidity. Validation on the newly curated, diverse PDI-Dataset demonstrates that the proposed method effectively uncovers geometric failure modes in state-of-the-art models—failures overlooked by conventional perceptual metrics—thereby offering a crucial evaluation tool for advancing physically plausible video generation and world model research.

3D structuregeometric consistencyquantitative evaluation

This work addresses the challenge that existing video generation methods struggle to model realistic three-dimensional physical motion due to their reliance on viewpoint-limited two-dimensional projections. To overcome this limitation, the authors propose a two-stage generative framework: first, Phys4View generates physically aware orthogonal foreground videos, which are then composited with background context by VideoSyn to produce complete scenes. Key innovations include a physics-aware attention mechanism, a geometry-enhanced cross-view attention module, and PhysMV—the first large-scale, multi-view synchronized video dataset comprising 160K sequences. Experimental results demonstrate that the proposed approach significantly outperforms current methods in terms of physical plausibility and spatiotemporal consistency.

3D motion dynamicsphysically consistent motionspatio-temporal coherence

Hot Scholars

SK

Seungryong Kim

Associate Professor, KAIST
Computer VisionMachine Learning
QZ

Qingwen Zhang

PhD Student, KTH (MPhil in HKUST)
autonomous drivingperceptionroboticsmapping
SL

Shuaicheng Liu

University of Electronic Science and Technology of China
Computer VisionComputational Photography
BZ

Bing Zeng

University of Electronic Science and Technology of China
Image and video processing