evaluate temporal consistency

Designs, implements, and analyzes metrics, benchmarks, and algorithms that measure, enforce, and optimize temporal and spatio‑temporal consistency in sequential data — e.g., preserving object identities, spatial relations, referential constraints, and stage‑wise predicates across frames. This includes building temporal consistency evaluation pipelines and benchmarks, constraint‑checking and optimization procedures, motion‑based smoothing and self‑supervised temporal regularizers, and methods for maintaining or enforcing cross‑frame coherence.

evaluatetemporalconsistency

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$213K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Current video large language models lack effective evaluation of temporal object consistency—such as identity preservation, state coherence, and cross-frame continuity—leading to an overestimation of their temporal reasoning capabilities. This work proposes the first object-trajectory-anchored benchmark for assessing temporal consistency, leveraging structured event timelines to focus on challenging scenarios involving occlusion, disappearance, and reappearance. We introduce a three-tier temporal necessity filtering protocol to ensure that questions strictly require ordered visual evidence, and we construct high-quality question-answer pairs through object tracking and human verification. The benchmark comprises 1,951 videos and 2,323 QA pairs, revealing significant deficiencies in mainstream models’ abilities to perform event counting, sequencing, and identity-sensitive reasoning.

object continuityobject identitytemporal object consistency

Existing image generation models struggle to capture the temporal evolution of the visual world, often failing to maintain cross-frame identity, spatial relationships, and causal ordering. To address this limitation, this work introduces ImageTime, a benchmark that evaluates models’ ability to generate temporally coherent images under sequential instructions using a four-keyframe protocol—comprising initial, action-onset, transition, and final states—thereby abstracting away video-level dynamics to focus on logical consistency between static frames. We propose a novel diagnostic evaluation framework centered on spatiotemporal consistency as a probing mechanism, featuring a hierarchical task structure and structured state predicates. Leveraging a VLM-as-judge paradigm enables interpretable scoring and failure attribution. Through multi-stage state definitions, temporal constraint modeling, and causal violation detection—augmented by GPT-5.5–driven automated assessment—our approach systematically uncovers the capability boundaries, failure modes, and concept drift phenomena of state-of-the-art models in temporal visual consistency.

coherent time imaginationimage generationspatiotemporal modeling

Existing video generation models often suffer from spatial distortions due to inconsistent 3D geometry, yet prevailing evaluation metrics such as FVD struggle to distinguish plausible motion dynamics from geometric inaccuracies. To address this gap, this work proposes the Spatial Geometric Consistency (SGC) metric, which segments static and dynamic regions in generated videos and enforces 3D consistency by partitioning the static background into spatially coherent subregions. For each subregion, SGC estimates depth and local camera poses, then quantifies geometric fidelity through multi-view pose divergence. SGC is the first metric capable of accurately detecting geometric distortions while preserving sensitivity to realistic dynamic content. It demonstrates superior robustness and discriminative power over existing methods on both real-world and synthetic videos, effectively addressing a critical limitation in current video generation evaluation frameworks.

3D spatial geometric consistencycamera pose divergencedynamic generated videos

ASurvey: Spatiotemporal Consistency in Video Generation

Feb 25, 2025
ZY
Zhiyu Yin
🏛️ Harbin Institute of Technology | Soochow University | Central South University | Peng Cheng Laboratory

This paper addresses the core challenge of spatiotemporal inconsistency in video generation—specifically, the lack of inter-frame motion coherence and spatial structural stability. We systematically survey technical approaches across five dimensions: foundational architectures, information representations, generative paradigms, post-processing techniques, and evaluation metrics—establishing, for the first time, a comprehensive taxonomy for spatiotemporal consistency in video generation. Our analysis reveals the intrinsic mechanisms by which diffusion models, Transformers, optical-flow guidance, temporal interpolation, and consistency regularization enable effective motion modeling and structural preservation. We unify multi-dimensional evaluation metrics—including TVD and FVD—into an extensible consistency assessment protocol. The study identifies key bottlenecks in current methods and outlines three critical future directions: controllable temporal modeling, implicit motion disentanglement, and standardized, unified evaluation benchmarks.

Addressing spatiotemporal consistency in video generationIdentifying gaps in high-quality video generation researchReviewing advances in video generation techniques

TCAFF: Temporal Consistency for Robot Frame Alignment

May 08, 2024
MB
Mason B. Peterson
🏛️ MIT

In global localization-denied environments (e.g., indoors or GPS-denied settings), multi-robot coordinate frames are difficult to align, leading to degraded collaborative performance. To address this, we propose the first multi-hypothesis coordinate frame alignment algorithm that requires no prior knowledge of initial robot poses. Our method jointly leverages sparse open-set semantic map matching, temporal consistency modeling, and robust geometric verification to achieve both initial alignment and long-term drift correction. Evaluated in a real-world scenario involving four robots collaboratively tracking six pedestrians, our approach achieves alignment accuracy approaching that of ground-truth localization systems (mean error < 0.15 m). The source code and hardware-collected dataset are publicly released. This work establishes a scalable, robust foundation for multi-robot cooperative perception and trajectory sharing—enabling precise, initialization-free coordination without external infrastructure.

Collaborative RobotsGPS Signal DegradationPosition Sharing

Latest Papers

What's happening recently
View more

This work addresses the susceptibility of video large language models to hallucinations in dynamic scenes, primarily due to insufficient explicit spatiotemporal modeling of object identities, states, and relationships over time. To mitigate this, the authors propose STEMO-Track, a novel framework that introduces explicit object trajectory modeling into video large language models for the first time. By integrating structured trajectory construction, chunked state extraction, and temporal aggregation mechanisms, STEMO-Track enables object-centric explicit spatiotemporal reasoning. Additionally, the authors introduce STEMO-Bench, the first human-verified benchmark specifically designed for evaluating object-centric factual consistency at a fine-grained level. Experimental results demonstrate that the proposed approach significantly reduces hallucination rates and outperforms state-of-the-art models in complex dynamic scenarios, achieving improved spatiotemporal reasoning consistency.

hallucinationmultimodal large language modelsobject tracking

It remains unclear whether current video prediction models genuinely understand the causal structure of the physical world or merely exploit superficial visual correlations. To address this, this work proposes CRONOS—a counterfactual evaluation benchmark based on interventions—introducing the first controllable manipulations of viewpoint, scene layout, object category, and appearance within a high-fidelity Unreal Engine environment. This framework establishes a reproducible standard for assessing physical consistency in video prediction. Experimental results demonstrate that state-of-the-art video generation models exhibit significant performance degradation under these interventions, revealing a fundamental deficiency in their capacity for true physical causal reasoning. These findings underscore the need for future models to incorporate explicit mechanisms for causal understanding of physical dynamics.

causal structurecounterfactual physical consistencyintervention-based benchmark

This work addresses geometric inconsistencies in text-to-video generation—such as object deformation, texture drift, and non-rigid background motion—by introducing a geometric consistency reward mechanism that explicitly optimizes temporal geometric structure during reinforcement fine-tuning of diffusion models. For the first time, geometric consistency is formulated as a directly optimizable objective without modifying the model’s latent space, making the approach applicable to complex dynamic scenes involving both camera and object motion. By integrating optical flow, depth-pose estimation, and feature correspondence techniques, the method effectively disentangles rigid background from dynamic object regions and evaluates their consistency separately. Experiments demonstrate that this approach substantially reduces temporal geometric artifacts while preserving high visual fidelity, outperforming strong existing baselines.

camera motiongeometric consistencyobject deformation

Hot Scholars

PW

Pengfei Wan

Head of Kling Video Generation Models, Kuaishou Technology
Generative ModelsComputer VisionMultimodal AIComputer Graphics
ZL

Ziwei Liu

Associate Professor, Nanyang Technological University
Computer VisionMachine LearningComputer Graphics
YS

Yiren Song

PH.D student, National University of Singapore
Generative AIDiffusionUnified model
QC

Qifeng Chen

HKUST
Computational PhotographyImage SynthesisGenerative AIAutonomous Driving
MH

Ming-Hsuan Yang

University of California at Merced; Google DeepMind
Computer VisionMachine LearningArtificial Intelligence