🤖 AI Summary
This study addresses the output homogenization in video generation models caused by repeated sampling, noting that existing global metrics fail to localize diversity collapse across spatiotemporal dimensions. We propose a dimension-level diagnostic framework that decouples diversity into six interpretable dimensions, quantified through computer vision pipelines and controlled prompt interventions. This approach enables a paradigm shift from merely measuring whether collapse occurs to diagnosing where and why it happens. Our analysis reveals that latent collapse in specific dimensions, such as motion and camera dynamics, stems from default mode convergence and implementation gaps, confirming this phenomenon as an inherent model capacity bottleneck rather than a consequence of prompt deviation. The evaluation framework has been open-sourced, offering a dimension-aware optimization pathway for controllable video generation.
📝 Abstract
Despite remarkable progress, video generation models often produce highly similar outputs when repeatedly sampled from the same prompt, limiting their usefulness for creative exploration. Existing diversity evaluations primarily rely on global scalar metrics, which obscure where diversity collapses in the spatiotemporal space of videos. We introduce DiVid, a dimension-level diagnostic framework that decomposes video generation diversity into six interpretable dimensions: Semantic, Style, Subject, Scene, Motion, and Camera. Each dimension is measured through a reproducible computer-vision pipeline and analyzed alongside quality and instruction faithfulness to examine potential trade-offs. Systematic evaluation of representative video generation models reveals that diversity is highly dimension-specific: models with strong global diversity scores still collapse on specific factors, particularly Motion and Camera. These rankings persist after filtering unfaithful generations, indicating genuine capability differences rather than off-prompt outputs. Beyond measurement, controlled prompt interventions identify two fundamental bottlenecks: default mode convergence, where models fall back to dominant patterns under open-ended prompts; and realization gaps, where models fail to faithfully realize diverse, explicitly requested alternatives, particularly for temporal factors. The larger faithfulness losses for temporal factors highlight the difficulty of controlling motion and camera variation through text alone. DiVid thus shifts the study of diversity from measuring whether it exists to diagnosing where and why it collapses, and provides actionable directions for dimension-aware training objectives and control signals. The framework will be released to facilitate future research on diverse and controllable video generation.