design perceptual metrics

Design and implement objective algorithms and scoring functions that estimate perceptual similarity or quality for signals and outputs across modalities (audio, speech, video, visual and vision–language), producing single-modality and multimodal evaluation metrics. Build end-to-end evaluation pipelines that compute and benchmark these metrics, run and analyze perceptual user studies, and validate or adapt metrics and perceptual features against human judgments, noisy or cross-domain data, and modality-specific characteristics.

designperceptualmetrics

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.34
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$215K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Modeling Beyond MOS: Quality Assessment Models Must Integrate Context, Reasoning, and Multimodality

May 26, 2025
MA
M. A. Kerkouri
🏛️ F-Initiatives | Université d’Orleans | Université Sorbonne Paris Nord | University of Dubai | IULM University

Current multimedia quality assessment models rely solely on Mean Opinion Scores (MOS) as supervision, neglecting semantic defects, user intent, and judgment rationale—resulting in poor interpretability and contextual adaptability. To address this, we propose a paradigm shift beyond scalar supervision, introducing the first systematic framework integrating context-aware modeling, evidence-driven reasoning, and cross-modal semantic alignment. Methodologically, we leverage Vision-Language Models (VLMs) to construct multimodal joint representations, incorporate contextual metadata modeling, and design an expert-level rationale generation mechanism guided by structured prompting and differentiable reasoning modules. Our evaluation framework emphasizes semantic alignment fidelity, reasoning faithfulness, and situational sensitivity. Experiments demonstrate substantial improvements in interpretability, task adaptability, and human alignment—advancing quality assessment from opaque prediction toward a trustworthy, robust, and human-AI collaborative paradigm.

Current benchmarks lack contextual metadata and expert rationalesMOS is insufficient for modern multimedia quality assessmentQuality models need context-awareness, reasoning, and multimodality

Do image and video quality metrics model low-level human vision?

Mar 20, 2025
DH
Dounia Hammou
🏛️ University of Cambridge | Netflix

This study investigates whether mainstream image/video quality metrics—such as SSIM, LPIPS, and VMAF—faithfully model human low-level visual mechanisms, including contrast sensitivity, contrast masking, and contrast matching. To this end, we introduce the first interpretability benchmark framework specifically designed for low-level vision properties, grounded in psychophysical principles: contrast sensitivity function testing, masking threshold estimation, and contrast-matching discrimination tasks. We systematically evaluate 33 full-reference metrics under this framework. Results show that LPIPS and MS-SSIM effectively capture contrast masking, whereas VMAF exhibits significant deficiencies; SSIM over-responds to high-frequency distortions, while MS-SSIM substantially mitigates this bias. Our analysis uncovers structural shortcomings in current metrics’ perceptual modeling, revealing misalignments with neurobiologically grounded vision principles. This work establishes a new paradigm for interpretable, neurophysiologically plausible quality assessment and provides empirical foundations for next-generation perceptual metrics.

Analyze strengths and weaknesses of 33 existing quality metrics.Evaluate image and video quality metrics' alignment with human vision.Test metrics' ability to model low-level human vision aspects.

Existing audio-visual generation models lack fine-grained, human-aligned automated evaluation metrics, and general-purpose multimodal models often fail to accurately capture human perceptual judgments. To address this gap, this work proposes the first human-centric benchmark for evaluating audio-visual generation, introducing ten fine-grained evaluation dimensions. The authors generate large-scale preference data through controlled perturbations and train a dedicated evaluator based on preference learning and multimodal consistency modeling. This evaluator produces continuous scores with calibrated confidence estimates, significantly improving alignment with human judgments. The resulting automated framework not only enables high-quality data filtering but also serves as a differentiable reward signal for human feedback in reinforcement learning, facilitating efficient and reliable assessment of generative models.

audio-video generationautomated evaluationevaluation benchmark

MAVERIX: Multimodal Audio-Visual Evaluation Reasoning IndeX

Mar 27, 2025
LX
Liuyue Xie
🏛️ Carnegie Mellon University | Amazon

Existing cross-modal evaluation frameworks lack standardized benchmarks for deep audio-visual fusion capabilities. To address this gap, we introduce AV-Bench—the first benchmark dedicated to audio-visual collaborative understanding—comprising 700 real-world videos and 2,556 questions requiring joint audio-visual reasoning across tightly coupled tasks, including synchrony assessment, causal inference, and event localization. We systematically define and evaluate models’ fine-grained audio-visual joint representation learning, encompassing both perceptual grounding and higher-order reasoning. Rigorous human annotation protocols and an open-source evaluation toolkit are established to ensure reproducibility and fairness. Evaluated on state-of-the-art models—including Gemini 1.5 Pro and o1—AV-Bench achieves ~70% accuracy, substantially outperforming prior benchmarks; human experts attain 95.1%, establishing a high-bar challenge. AV-Bench thus fills a critical void in cross-modal perception evaluation and sets a new standard for multimodal intelligence assessment.

Lacks standardized evaluation for audio-visual cross-modality modelsNeeds benchmark to assess video-audio integration in AIRequires framework mimicking human multimodal perception tasks

AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Oct 23, 2024
KS
Kim Sung-Bin
🏛️ POSTECH | KAIST | Yonsei University

Audio-visual large language models (AV-LLMs) suffer from cross-modal hallucination—erroneous associations between audio and visual signals—yet lack dedicated, standardized evaluation benchmarks. Method: We introduce AVHBench, the first benchmark specifically designed to evaluate cross-modal hallucination in AV-LLMs. It formally defines and quantifies such hallucinations, establishes a three-dimensional evaluation framework covering perception, alignment matching, and multimodal reasoning, and constructs a test set via a synergistic strategy combining multi-granularity aligned samples with human annotation and adversarial perturbation to enable fine-grained attribution analysis. Results: Experiments reveal that state-of-the-art AV-LLMs are consistently vulnerable to modality-crossing interference, inducing widespread hallucination. Crucially, fine-tuning solely on AVHBench significantly enhances hallucination robustness. This work provides foundational tools and a methodological framework for trustworthy evaluation and optimization of audio-visual multimodal models.

Address hallucinations caused by audio-visual signal misinterpretations.Develop a benchmark to improve robustness against hallucinations.Evaluate audio-visual LLMs' cross-modal perception and comprehension.

Latest Papers

What's happening recently
View more

This study investigates whether vision-language models (VLMs) can effectively emulate human judgments in perceptual image quality assessment, potentially replacing costly psychophysical experiments. For the first time, we systematically evaluate six VLMs—four closed-source and two open-source—against human judgments across three dimensions: contrast, color saturation, and overall preference, using established psychophysical data as a benchmark. Our analysis integrates attribute-weighted evaluation and intra-model consistency metrics. Results show that VLMs achieve up to 0.93 correlation with human judgments on color saturation but perform notably weaker on contrast. Most models align with human behavior by prioritizing color saturation in overall preference. A key contribution is revealing the trade-off between model self-consistency and human alignment, and demonstrating that enhancing perceptual separability improves human-model agreement.

Attribute DependencyHuman AlignmentPerceptual Image Quality Assessment

This study addresses the limitations of existing music–flavor crossmodal research, which has been constrained by small-scale, high-cost perceptual data. The authors overcome this bottleneck through two complementary experiments: first, they demonstrate that crossmodal structures identified in small-scale human-annotated data generalize to large-scale synthetically labeled audio; second, they evaluate the alignment between chemically derived computational flavor profiles and human perception via an online auditory experiment. Results show that crossmodal structures remain stable across different supervision regimes, and computational flavor representations exhibit strong agreement with human ratings (p<0.0001, Mantel r=0.45, Procrustes m²=0.51). This work provides the first evidence that synthetically labeled data can preserve genuine perceptual structure and establishes a reproducible framework for computational flavor modeling and validation, releasing both dataset and code.

cross-modal alignmentmultimodal datasetmusic-taste correspondence

Existing music understanding benchmarks suffer from limitations such as static evaluation protocols, high resource demands, absence of authentic perceptual assessment, and lack of cross-modal comparability. This work proposes MusICA-MetaBench, a novel framework that introduces a pedagogy-aligned, on-demand generation paradigm for evaluating multimodal large language models’ musical perception capabilities. Leveraging symbolic inputs like MusicXML provided by users, the framework automatically constructs multimodal multiple-choice questions encompassing audio, sheet music images, and symbolic representations. It integrates structured musical representations, predefined question templates, and statistical reliability analysis to ensure validity. Evaluation on the ChoraleBricks dataset demonstrates the effectiveness of the generated questions and identifies the minimal benchmark size required to support statistically significant model comparisons.

cross-modal evaluationlarge language modelsmultimodal benchmarking

Existing video understanding benchmarks are largely confined to single videos or static images, making them inadequate for evaluating models’ capacity to comprehend complex cross-temporal and cross-view interactions across multiple videos. To address this gap, this work introduces the first comprehensive benchmark for multimodal multi-video perception, encompassing 14 subtasks and 5,000 structured question-answer pairs. The benchmark integrates both existing datasets and newly annotated video content, covering a diverse range of visual scenarios. Experimental results demonstrate that current state-of-the-art multimodal large language models exhibit significant performance degradation when processing multi-video inputs, revealing critical limitations in their ability to perform coordinated understanding across multiple video streams. These findings underscore the necessity and challenge of the proposed benchmark in advancing multi-video reasoning capabilities.

evaluation benchmarkmulti-modal video understandingmulti-video interaction

Hot Scholars

GZ

Guangtao Zhai

Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI Evaluation
HD

Huiyu Duan

Shanghai Jiao Tong University
Multimedia Signal Processing
PW

Pengfei Wan

Head of Kling Video Generation Models, Kuaishou Technology
Generative ModelsComputer VisionMultimodal AIComputer Graphics
YZ

Yucheng Zhu

Shanghai Jiaotong University
Multimedia Signal Processing
LS

Li Song

Professor of Electronic Engineering, Shanghai Jiao Tong University
Video CodingImage ProcessingComputer Vision