Score
Designs and implements evaluation metrics, benchmarks, and analysis pipelines that assess perceptual and technical quality of video content across multiple spatial and temporal scales. This includes constructing and validating local (frame- or patch-level), cross-event/temporal-context, and global (entire-sequence) evaluation protocols and quantifying the sensitivity or reasoning complexity of models at each level.
Existing large multimodal models (LMMs) lack systematic evaluation for video quality understanding. Method: We introduce VQ-Bench, the first dedicated multimodal benchmark for this task, covering three video sources—natural, AI-generated content (AIGC), and computer graphics (CG)—and four question types: Yes/No, What-How, open-ended QA, and pairwise quality comparison. It is the first benchmark to incorporate AIGC-specific distortion dimensions. We propose a systematic evaluation framework featuring a multi-granularity QA design and cross-source sampling, validated via expert annotation to yield 2,378 high-quality QA pairs. Contribution/Results: Comprehensive evaluation across 17 state-of-the-art LMMs reveals a substantial performance gap between model and human capabilities in video quality understanding. VQ-Bench establishes the first reproducible benchmark for this task and identifies concrete directions for future improvement.
Existing video evaluation metrics are primarily designed for short videos and struggle to effectively capture long-range characteristics such as narrative richness and global causal consistency in long-form content. This work proposes decoupling long-video evaluation from short-video assessment by explicitly modeling long-range context as an orthogonal dimension to short-term visual quality. To this end, the authors introduce a dedicated evaluation framework for structural consistency in long videos, integrating shot-level dynamics modeling, perturbation-based stress tests—including shot shuffling and narrative disruption—and a human-annotated dataset of long-range features, Long-CODE. The proposed metric achieves state-of-the-art correlation with human judgments and complements existing short-video evaluation benchmarks, together forming a comprehensive assessment suite for video generation.
Current evaluation methods for video generation are largely confined to basic prompt adherence—assessing whether outputs are “correct”—while neglecting critical dimensions of cinematic excellence such as aesthetic quality, performative expressiveness, and audiovisual coherence that determine whether results are truly “excellent.” Moreover, existing automated metrics lack professional credibility. To bridge this gap, this work proposes EvalVerse, the first framework to formalize expert knowledge from professional film production into a structured evaluation taxonomy. Leveraging large-scale expert-annotated data, EvalVerse employs expert-calibrated fine-tuning of vision-language models combined with chain-of-thought reasoning to enable a qualitative leap from correctness to excellence. The framework supports fine-grained diagnostics for multi-shot sequences and complex audiovisual tasks, significantly enhancing assessment fidelity for professional-grade video quality while remaining compatible with conventional metrics, thereby establishing a reliable infrastructure for reward modeling and intelligent evaluation.
Existing video generation evaluation benchmarks suffer from two key limitations: conventional multi-metric embedding approaches exhibit low correlation with human preferences, while LLM-based methods, though reasoning-capable, lack deep understanding of video quality and cross-modal consistency. To address this, we introduce the first comprehensive, human-preference-oriented video generation benchmark, pioneering the integration of multimodal large language models (MLLMs) across the entire evaluation domain. We propose a novel few-shot scoring paradigm coupled with a chain-of-query mechanism to enable scalable, structured, and holistic automated assessment across multiple dimensions—including temporal coherence, visual fidelity, and text-video alignment. Empirical validation on state-of-the-art models (e.g., Sora) demonstrates significantly improved human alignment across all dimensions compared to existing benchmarks, greater objectivity in ambiguous cases, and emergent capability surpassing human judgment in certain scenarios.
This work addresses online video quality assessment (VQA), targeting the fundamental limits of spatiotemporal redundancy compression to achieve optimal trade-offs between efficiency and accuracy. We propose a joint spatiotemporal sampling strategy that reduces spatial resolution and frame rate to ≤10% of the original while incurring <8% performance degradation. The method comprises a lightweight spatial feature extractor, an efficient temporal fusion module, and a global quality regression network—forming a low-latency, real-time-capable VQA architecture. To our knowledge, this is the first study to systematically characterize the performance tolerance thresholds for spatiotemporal compression in VQA. Extensive experiments on six mainstream public benchmarks demonstrate strong generalization and robustness. Our approach establishes the first practical, high-accuracy, low-overhead solution for real-time VQA in edge-computing and streaming scenarios.
Existing video generation evaluation benchmarks primarily target photorealistic content and are ill-suited for assessing character-centric animation, particularly in aspects such as character consistency, exaggerated motion, and stylized expression, while also lacking support for open-domain content and customizable evaluation. This work proposes the first image-to-video generation benchmark specifically designed for character-focused animation, systematically formalizing the Twelve Principles of Animation into quantifiable dimensions. It establishes a multidimensional evaluation framework incorporating metrics like intellectual property fidelity, semantic consistency, and motion plausibility. Leveraging vision-language models, the benchmark enables fine-grained, scalable automated scoring that unifies closed-set standardized assessment with open-set flexible analysis. Experiments demonstrate strong alignment with human judgment and reveal distinctive shortcomings of current image-to-video models in animation generation, significantly outperforming realism-oriented benchmarks in both discriminative power and informativeness.
Existing no-reference video quality assessment (NR-VQA) methods are designed for camera-captured videos and suffer significant performance degradation on rendered videos (e.g., gaming, VR), primarily due to their neglect of temporal artifacts. This work presents the first systematic study of NR-VQA for rendered content. We introduce RenderVQA—the first large-scale, multi-scenario, multi-rendering-setup dataset featuring subjective quality scores across diverse display types. To address the unique distortions introduced by temporal super-resolution and frame generation, we propose RQNet, a deep learning–based metric that jointly models spatial fidelity and temporal stability, explicitly capturing time-domain degradations. Experiments demonstrate that RQNet substantially outperforms state-of-the-art NR-VQA methods on RenderVQA (average PLCC improvement of 0.21). Moreover, it enables reliable benchmarking of super-resolution techniques and quantitative evaluation of frame-generation strategies, establishing a robust, deployable tool for real-time rendering quality analysis.
Existing vision-language models lack effective evaluation for long-form video quality understanding, as prevailing benchmarks are confined to short clips and isolated distortions, overlooking temporal continuity and complex reasoning. To address this gap, this work proposes LongVQUBench—the first hierarchical evaluation benchmark tailored for long video quality understanding—comprising over 1,200 diverse real-world long videos and 1,500 hierarchically structured question-answer pairs that span local event perception, cross-event reasoning, and global quality judgment. It further introduces a novel “needle-in-a-haystack distortion” questioning paradigm to probe fine-grained recognition capabilities. Systematic evaluation of 14 state-of-the-art models reveals a marked performance decline with increasing video length and reasoning depth, exposing critical limitations in long-range temporal integration and perceptual attribution.
This study addresses the temporal inconsistency—such as flickering and motion jerkiness—that commonly arises in video compression at low bitrates, a phenomenon whose relationship with compression intensity remains poorly understood. The authors systematically evaluate the impact of mainstream codecs (AV1, HEVC, VP9, and H.264) on inter-frame consistency across varying bitrates and content types, employing objective metrics to quantify temporal distortion. Their findings reveal that temporal consistency degrades nonlinearly with increasing compression strength and that videos with unpredictable dynamics exhibit greater temporal instability than those with high but predictable motion—challenging the conventional assumption that motion magnitude alone dictates encoding difficulty. These results underscore the urgent need to incorporate temporally aware quality metrics into existing compression pipelines to enhance perceptual visual fidelity.
This work addresses the lack of systematic evaluation of large multimodal language models in video aesthetic perception, particularly their limited exploration in fundamental human aesthetic quality assessment tasks. To bridge this gap, we introduce VideoAesBench—the first comprehensive benchmark dedicated to video aesthetics—comprising 1,804 videos sourced from diverse domains including user-generated content (UGC), AI-generated content (AIGC), compression artifacts, robotics, and gaming. The benchmark features a structured annotation framework spanning three dimensions: visual form, style, and emotion, and incorporates diverse question formats such as open-ended descriptions. Systematic evaluation of 23 leading large models reveals that current approaches possess only rudimentary aesthetic perception capabilities, exhibiting incomplete coverage and limited accuracy, thereby offering a critical reference for future research on interpretable video aesthetics.