🤖 AI Summary
Current AI-generated human-centric videos often suffer from quality artifacts and semantic inconsistencies, and lack comprehensive multidimensional evaluation benchmarks. To address this gap, this work introduces HVEval+, the first large-scale human-annotated dataset encompassing three critical dimensions: spatial quality, temporal coherence, and text-video alignment. Furthermore, we propose MoE-Rater, a unified multitask evaluation model that integrates Mixture-of-Projection Experts (MoPE) and Mixture-of-LoRA Experts (MoLE), trained via a three-stage strategy. Evaluated on both HVEval+ and Human-AGVQA, MoE-Rater significantly outperforms existing methods, offering a reliable and holistic assessment tool to advance text-to-video generation models.
📝 Abstract
AI-generated human-centric videos play a crucial role in a wide range of modern applications. However, they often suffer from quality issues and semantic mismatches, underscoring the importance of effective quality assessment for such videos. To this end, we extend our previous dataset HVEval with pairwise preference annotations, resulting in HVEval+, the largest holistic quality assessment dataset for AI-generated human-centric videos, which comprises 1k prompts based on a comprehensive taxonomy, 20k videos generated by 24 text-to-video (T2V) models, and extensive human annotations, including 60k mean opinion scores (MOSs) and 60k preference pairs across 3 dimensions (i.e., spatial quality, temporal quality, and text-video correspondence), as well as 20k category-specific question-answer (Q&A) pairs. Along with the HVEval+ dataset, we further propose MoE-Rater, a Mixture-of-Experts (MoE)-inspired and multimodal large language model (MLLM)-based all-in-one method that supports multi-dimensional quality rating, multi-dimensional pairwise comparison, and category-specific question answering within a single model. Specifically, we introduce Mixture of Projector Experts (MoPE) and Mixture of LoRA Experts (MoLE), together with a three-stage training strategy consisting of task-aware pre-training, task-specific adaptation, and adaptive routing optimization, to effectively unify multiple tasks, resulting in superior performance on both HVEval+ and Human-AGVQA datasets. Extensive experiments and comprehensive analysis demonstrate the significant potential of the HVEval+ dataset and the MoE-Rater method in advancing AI-generated video quality assessment and further facilitating the evaluation and optimization of T2V models.