Score
Design and implement objective algorithms and scoring functions that estimate perceptual similarity or quality for signals and outputs across modalities (audio, speech, video, visual and vision–language), producing single-modality and multimodal evaluation metrics. Build end-to-end evaluation pipelines that compute and benchmark these metrics, run and analyze perceptual user studies, and validate or adapt metrics and perceptual features against human judgments, noisy or cross-domain data, and modality-specific characteristics.
Current multimedia quality assessment models rely solely on Mean Opinion Scores (MOS) as supervision, neglecting semantic defects, user intent, and judgment rationale—resulting in poor interpretability and contextual adaptability. To address this, we propose a paradigm shift beyond scalar supervision, introducing the first systematic framework integrating context-aware modeling, evidence-driven reasoning, and cross-modal semantic alignment. Methodologically, we leverage Vision-Language Models (VLMs) to construct multimodal joint representations, incorporate contextual metadata modeling, and design an expert-level rationale generation mechanism guided by structured prompting and differentiable reasoning modules. Our evaluation framework emphasizes semantic alignment fidelity, reasoning faithfulness, and situational sensitivity. Experiments demonstrate substantial improvements in interpretability, task adaptability, and human alignment—advancing quality assessment from opaque prediction toward a trustworthy, robust, and human-AI collaborative paradigm.
This study investigates whether mainstream image/video quality metrics—such as SSIM, LPIPS, and VMAF—faithfully model human low-level visual mechanisms, including contrast sensitivity, contrast masking, and contrast matching. To this end, we introduce the first interpretability benchmark framework specifically designed for low-level vision properties, grounded in psychophysical principles: contrast sensitivity function testing, masking threshold estimation, and contrast-matching discrimination tasks. We systematically evaluate 33 full-reference metrics under this framework. Results show that LPIPS and MS-SSIM effectively capture contrast masking, whereas VMAF exhibits significant deficiencies; SSIM over-responds to high-frequency distortions, while MS-SSIM substantially mitigates this bias. Our analysis uncovers structural shortcomings in current metrics’ perceptual modeling, revealing misalignments with neurobiologically grounded vision principles. This work establishes a new paradigm for interpretable, neurophysiologically plausible quality assessment and provides empirical foundations for next-generation perceptual metrics.
Existing audio-visual generation models lack fine-grained, human-aligned automated evaluation metrics, and general-purpose multimodal models often fail to accurately capture human perceptual judgments. To address this gap, this work proposes the first human-centric benchmark for evaluating audio-visual generation, introducing ten fine-grained evaluation dimensions. The authors generate large-scale preference data through controlled perturbations and train a dedicated evaluator based on preference learning and multimodal consistency modeling. This evaluator produces continuous scores with calibrated confidence estimates, significantly improving alignment with human judgments. The resulting automated framework not only enables high-quality data filtering but also serves as a differentiable reward signal for human feedback in reinforcement learning, facilitating efficient and reliable assessment of generative models.
Existing cross-modal evaluation frameworks lack standardized benchmarks for deep audio-visual fusion capabilities. To address this gap, we introduce AV-Bench—the first benchmark dedicated to audio-visual collaborative understanding—comprising 700 real-world videos and 2,556 questions requiring joint audio-visual reasoning across tightly coupled tasks, including synchrony assessment, causal inference, and event localization. We systematically define and evaluate models’ fine-grained audio-visual joint representation learning, encompassing both perceptual grounding and higher-order reasoning. Rigorous human annotation protocols and an open-source evaluation toolkit are established to ensure reproducibility and fairness. Evaluated on state-of-the-art models—including Gemini 1.5 Pro and o1—AV-Bench achieves ~70% accuracy, substantially outperforming prior benchmarks; human experts attain 95.1%, establishing a high-bar challenge. AV-Bench thus fills a critical void in cross-modal perception evaluation and sets a new standard for multimodal intelligence assessment.
Audio-visual large language models (AV-LLMs) suffer from cross-modal hallucination—erroneous associations between audio and visual signals—yet lack dedicated, standardized evaluation benchmarks. Method: We introduce AVHBench, the first benchmark specifically designed to evaluate cross-modal hallucination in AV-LLMs. It formally defines and quantifies such hallucinations, establishes a three-dimensional evaluation framework covering perception, alignment matching, and multimodal reasoning, and constructs a test set via a synergistic strategy combining multi-granularity aligned samples with human annotation and adversarial perturbation to enable fine-grained attribution analysis. Results: Experiments reveal that state-of-the-art AV-LLMs are consistently vulnerable to modality-crossing interference, inducing widespread hallucination. Crucially, fine-tuning solely on AVHBench significantly enhances hallucination robustness. This work provides foundational tools and a methodological framework for trustworthy evaluation and optimization of audio-visual multimodal models.
This study investigates whether vision-language models (VLMs) can effectively emulate human judgments in perceptual image quality assessment, potentially replacing costly psychophysical experiments. For the first time, we systematically evaluate six VLMs—four closed-source and two open-source—against human judgments across three dimensions: contrast, color saturation, and overall preference, using established psychophysical data as a benchmark. Our analysis integrates attribute-weighted evaluation and intra-model consistency metrics. Results show that VLMs achieve up to 0.93 correlation with human judgments on color saturation but perform notably weaker on contrast. Most models align with human behavior by prioritizing color saturation in overall preference. A key contribution is revealing the trade-off between model self-consistency and human alignment, and demonstrating that enhancing perceptual separability improves human-model agreement.
This study addresses the limitations of existing music–flavor crossmodal research, which has been constrained by small-scale, high-cost perceptual data. The authors overcome this bottleneck through two complementary experiments: first, they demonstrate that crossmodal structures identified in small-scale human-annotated data generalize to large-scale synthetically labeled audio; second, they evaluate the alignment between chemically derived computational flavor profiles and human perception via an online auditory experiment. Results show that crossmodal structures remain stable across different supervision regimes, and computational flavor representations exhibit strong agreement with human ratings (p<0.0001, Mantel r=0.45, Procrustes m²=0.51). This work provides the first evidence that synthetically labeled data can preserve genuine perceptual structure and establishes a reproducible framework for computational flavor modeling and validation, releasing both dataset and code.
Existing music understanding benchmarks suffer from limitations such as static evaluation protocols, high resource demands, absence of authentic perceptual assessment, and lack of cross-modal comparability. This work proposes MusICA-MetaBench, a novel framework that introduces a pedagogy-aligned, on-demand generation paradigm for evaluating multimodal large language models’ musical perception capabilities. Leveraging symbolic inputs like MusicXML provided by users, the framework automatically constructs multimodal multiple-choice questions encompassing audio, sheet music images, and symbolic representations. It integrates structured musical representations, predefined question templates, and statistical reliability analysis to ensure validity. Evaluation on the ChoraleBricks dataset demonstrates the effectiveness of the generated questions and identifies the minimal benchmark size required to support statistically significant model comparisons.
为了解决音频视频编辑中模态选择性问题,AVENUE提出了包含多样化编辑类型的数据集和模态感知的评估框架。
Existing video understanding benchmarks are largely confined to single videos or static images, making them inadequate for evaluating models’ capacity to comprehend complex cross-temporal and cross-view interactions across multiple videos. To address this gap, this work introduces the first comprehensive benchmark for multimodal multi-video perception, encompassing 14 subtasks and 5,000 structured question-answer pairs. The benchmark integrates both existing datasets and newly annotated video content, covering a diverse range of visual scenarios. Experimental results demonstrate that current state-of-the-art multimodal large language models exhibit significant performance degradation when processing multi-video inputs, revealing critical limitations in their ability to perform coordinated understanding across multiple video streams. These findings underscore the necessity and challenge of the proposed benchmark in advancing multi-video reasoning capabilities.