🤖 AI Summary
Existing video generation evaluation methods predominantly focus on aesthetics or text alignment, making it difficult to quantify object motion defects. This work addresses this limitation by constructing the VidMotion dataset and proposing MotionInsight, a diagnostic framework that shifts the evaluation paradigm from implicit RGB frame observation to explicit motion-space diagnosis. By integrating motion-aware representations, Group Relative Policy Optimization (GRPO) reinforcement learning, and motion-specific reward mechanisms, the proposed approach enables fine-grained assessment of object consistency, temporal continuity, and physical plausibility. Furthermore, MotionInsight yields three-dimensional scores aligned with human preferences and provides fact-based interpretability analysis for identifying specific motion artifacts, offering a more rigorous and transparent methodology for evaluating generated videos.
📝 Abstract
Despite rapid progress in video generation models, they still exhibit obvious motion deficiencies, often manifested as incorrect object motion. However, most existing video quality evaluations focus on aesthetic quality or text-video alignment. To address this gap, we study object-centric motion fidelity assessment, evaluating target objects along object consistency, motion continuity, and physical plausibility. To achieve this, we first introduce VidMotion, a diagnostic dataset of 6,879 videos with designated moving objects and fine-grained annotations including dimension-wise scores and failure causes. We further propose MotionInsight, a diagnostic evaluator that shifts assessment from implicit RGB-frame observation to explicit motion-space diagnosis. By constructing motion-aware representations, MotionInsight makes subtle motion deficiencies more observable. We also introduce motion-specific rewards during GRPO to transform observed motion into a diagnostic assessment. Experiments demonstrate that MotionInsight provides an effective basis for diagnosing object motion deficiencies, producing human-aligned scores along three dimensions and grounded explanations.