Towards Subject Consistency over Dynamic Subject Sets in Video Generation

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of consistency evaluation for dynamic subject sets and the failure of conventional metrics in long video generation. To this end, it proposes the DynSC-Eval framework and a DiffusionNFT post-training method. The former constructs an object-level, multi-granularity measurement paradigm by tracking subject lifecycles, with its sensitivity validated through synthetic experiments. The latter optimizes generation consistency by integrating diffusion model reinforcement learning with curriculum learning strategies. Experimental results demonstrate that the proposed approach reduces inconsistency metrics by 13.82% and 5.66% for the Wan and SANA models, respectively, while effectively preserving their long-duration generation capabilities.
📝 Abstract
We argue that as video generation extends to longer durations, subject consistency should be evaluated over \textit{dynamic subject sets}. We therefore introduce \textbf{DynSC-Eval}, an evaluation framework that dynamically tracks eligible subjects throughout their visible lifespans and measures local continuity and global identity preservation using six complementary object-level metrics, with explicit detection of inconsistency events. To validate its effectiveness, we design synthetic experiments that actively inject inconsistency events, demonstrating both the sensitivity of DynSC-Eval and the limitations of existing metrics. Evaluations of diverse models on 5s, 15s, and 60s video generation further reveal substantial subject consistency differences that are obscured by conventional metrics. Beyond evaluation, we construct rewards from DynSC-Eval and apply DiffusionNFT post-training in an autonomous-driving testbed. On 5s generation, our approach reduces the six inconsistency metrics by an average of 13.82\% for Wan-2.1-1.3B and 5.66\% for SANA-2B, with improvements also observed on the I2V model ReSim. Qualitative comparisons further demonstrate the effectiveness of our method. We then extend generation to 10s and 30s through curriculum learning and show that consistency optimization remains effective while largely preserving other capabilities.
Problem

Research questions and friction points this paper is trying to address.

video generation
subject consistency
dynamic subject sets
evaluation metrics
long-duration video
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dynamic Subject Sets
Video Generation Evaluation
Subject Consistency
Reward-based Post-training
Curriculum Learning