listening test design

Designing and running human perceptual evaluations for audio systems, including stimulus selection, protocol design, and statistical analysis to compare timbre, utility, and baselines. Used to validate trade-offs where automatic metrics diverge from human judgments and to measure improvements on uncommon or adversarial prompts.

listeningtestdesign

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Exploring Perceptual Audio Quality Measurement on Stereo Processing Using the Open Dataset of Audio Quality

Dec 11, 2025
PM
Pablo M. Delgado
🏛️ Fraunhofer Institute for Integrated Circuits IIS | Ball State University | Netflix, Inc.

This study investigates how stereo processing—specifically Mid/Side versus Left/Right encoding—affects subjective audio quality perception, and evaluates the predictive accuracy of mainstream objective metrics (e.g., PEAQ, ITU-R BS.1387, DNSMOS) under spatial distortions. Leveraging the ODAQ stereo extension dataset with corresponding Mean Opinion Scores (MOS), we conduct time–frequency domain metric comparisons and statistical modeling. Our analysis quantitatively reveals, for the first time, the critical interplay between bottom-up auditory mechanisms and top-down contextual factors in stereo quality prediction. Results show that timbre-oriented metrics remain robust under simple distortions but degrade significantly under spatial distortions; current models exhibit systematic bias due to their neglect of spatial dimensions. We propose a novel three-dimensional perceptual evaluation paradigm integrating temporal, spectral, and spatial cues—providing both theoretical foundation and methodological support for next-generation audio quality metrics.

Evaluating stereo processing impact on audio qualityImproving models for timbral and spatial quality perceptionTesting objective metrics with subjective ratings dataset

Evaluation of Audio Compression Codecs

Nov 14, 2025
TT
Thien T. Duong

This study addresses the challenge of balancing compression efficiency and perceptual audio quality in audio codec selection. Methodologically, it introduces a human auditory perception–centered evaluation framework integrating the Perceptual Evaluation of Audio Quality (PEAQ) objective model, multi-bitrate encoding performance testing, time-frequency spectrogram visualization, and multidimensional quality analysis to quantitatively characterize distortion mechanisms affecting perceived sound quality. Its key contribution lies in the first systematic, cross-codec comparison—under standardized experimental conditions—of mainstream codecs (e.g., MP3, AAC, Opus, FLAC) along their rate–perceptual-quality trade-off curves, revealing distinct patterns of perceptual degradation. The results provide reproducible empirical evidence and application-oriented, quality-efficiency co-optimization guidelines for codec selection across diverse use cases.

Analyzing how compression affects human-perceived sound fidelityEvaluating audio codecs' compression efficiency and perceptual qualityProviding selection guidance for digital audio compression schemes

This study addresses the misalignment between human subjective preferences and objective evaluation metrics in music generation, focusing on text-audio alignment and musical quality. We construct a large-scale benchmark comprising 6,000 generated tracks from 12 state-of-the-art models and conduct 15,000 pairwise auditory comparison trials across 2,500 human participants—the first such large-scale human preference study in the domain. Through rigorous statistical correlation analysis, we systematically quantify the consistency between automated metrics (e.g., CLAP, FAD) and human judgments, revealing substantial discrepancies between existing metrics and true perceptual preferences. Our work delivers the most comprehensive ranking of both models and evaluation metrics to date, and publicly releases a high-quality human preference dataset. This resource establishes a robust, human-centered benchmark to drive paradigm shifts in evaluation methodology and model optimization for text-to-music generation.

Assessing text-audio alignment and music quality metricsEvaluating music generation models via human preferencesRanking state-of-the-art models using human survey data

This work addresses the challenge of automated audio aesthetic quality assessment by proposing the first unified no-reference cross-domain (speech/music/environmental sound) aesthetic scoring framework. Methodologically, it introduces a novel four-dimensional perceptual disentanglement annotation scheme and constructs a single-sample prediction model based on self-supervised representations and multi-task regression, integrating perception-driven feature extraction with culture-robust aesthetic modeling to enable fine-grained, interpretable aesthetic quantification. Evaluated across diverse audio domains, the predicted scores achieve high correlation with human Mean Opinion Scores (MOS), yielding Spearman’s ρ > 0.89—significantly outperforming existing approaches. To foster reproducibility and downstream applications, the project open-sources the trained models, annotated dataset, and implementation code, providing a reliable tool for audio data filtering, pseudo-label generation, and generative audio evaluation.

Automated prediction of audio aestheticsEnhances quality assessment in audio processingReduces reliance on human evaluation

Preference-Based Learning in Audio Applications: A Systematic Analysis

Nov 17, 2025
AB
Aaron Broukhim
🏛️ UC San Diego

Audio preference learning remains severely underexplored, with no systematic literature review, standardized benchmarks, or explicit modeling of temporal dynamics. Method: Following the PRISMA framework, we conduct the first systematic review revealing that only 6% of audio studies employ preference learning; we propose a novel multi-source preference signal integration paradigm—combining synthetic, automated, and human feedback—and design a multi-stage training framework targeting subjective dimensions (e.g., naturalness, musicality). We validate our approach using rankSVM and RLHF, analyzing alignment between objective metrics and human judgments. Contribution/Results: We find poor correlation (mean <0.3) between conventional objective metrics and human preferences, and demonstrate that explicit temporal modeling significantly improves preference consistency. Our work establishes foundational resources—including high-quality audio preference datasets and evaluation benchmarks—thereby enabling more reliable, human-aligned assessment of generative audio models.

Audio preference learning lacks standardized benchmarks and systematic investigation of temporal factorsPreference learning is significantly underexplored in audio applications compared to text domainsTraditional audio metrics often misalign with human judgments across different evaluation contexts

Latest Papers

What's happening recently
View more

This work addresses the concern that high performance of current audio-language models on standard benchmarks may stem from reliance on textual priors rather than genuine utilization of acoustic signals for auditory comprehension. To investigate this, the authors propose a diagnostic framework that systematically quantifies the extent to which models actually depend on audio inputs. Through ablation studies and segment-level audio analysis, they evaluate eight prominent audio-language models across three major benchmarks. Their findings reveal that models retain 60–72% of their original accuracy even without any audio input, and only 3.0–4.2% of questions genuinely require the full audio signal—most tasks can be solved using brief local segments. These results challenge prevailing evaluation paradigms and expose a critical limitation: existing benchmarks inadequately assess true auditory reasoning capabilities.

audio relianceaudio-language modelsauditory understanding

This work addresses the limitations of existing audio separation evaluation metrics, which often misalign with human perception, offer coarse granularity, and rely on ground-truth reference signals, while subjective listening tests remain costly and non-scalable. To overcome these challenges, we propose SAM Audio Judge—a reference-free, multimodal, fine-grained objective evaluation framework that, for the first time, unifies assessment across speech, music, and general sound events. Our approach integrates textual, visual, and temporal span prompts to achieve perceptually aligned scoring through an end-to-end trainable multimodal architecture. It delivers fine-grained evaluations along four dimensions: recall, precision, fidelity, and overall quality, significantly improving correlation with human judgments. Furthermore, the method demonstrates broad applicability in data curation, pseudo-label generation, and model re-ranking. Code and pretrained models are publicly released.

audio separationmultimodal evaluationobjective metric

This study addresses the limited capability of large audio language models in fundamental auditory perception—such as pitch, loudness, and spatial location—where their performance often approaches random guessing and fails to match human superiority in comparative tasks. To systematically evaluate this gap, the authors propose SonicBench, the first psychophysics-based benchmark integrating controllable audio generation, a dual-task paradigm of identification and comparison, linear probing analysis, and controlled experiments. Findings reveal that frozen audio encoders already capture physical auditory cues effectively (accuracy ≥60%), indicating that the primary bottleneck lies not in perceptual encoding but in subsequent alignment and decoding stages. This work provides the first systematic characterization of the boundaries of audio language models in basic auditory perception.

auditory understandingLarge Audio Language Modelsperceptual dimensions

Existing audio-visual generation models lack fine-grained, human-aligned automated evaluation metrics, and general-purpose multimodal models often fail to accurately capture human perceptual judgments. To address this gap, this work proposes the first human-centric benchmark for evaluating audio-visual generation, introducing ten fine-grained evaluation dimensions. The authors generate large-scale preference data through controlled perturbations and train a dedicated evaluator based on preference learning and multimodal consistency modeling. This evaluator produces continuous scores with calibrated confidence estimates, significantly improving alignment with human judgments. The resulting automated framework not only enables high-quality data filtering but also serves as a differentiable reward signal for human feedback in reinforcement learning, facilitating efficient and reliable assessment of generative models.

audio-video generationautomated evaluationevaluation benchmark

This study addresses a critical yet overlooked issue in evaluating large audio language models: their potential reliance on cues from evaluation protocols—such as provided labels or reference information—rather than genuine comprehension of audio content, which can artificially inflate human-model agreement. To probe such protocol-level shortcuts, the authors introduce a novel auditing framework employing three adversarial manipulations: feature blueprint substitution, reference information perturbation, and option order swapping. Applying this approach across six prominent models and four speech attributes, the experiments reveal substantial shortcut dependence; for instance, emotion recognition accuracy drops below 0.10 for several models, and Qwen3-Omni-Thinking consistently selects the same position in A/B tests regardless of content. These findings underscore the necessity of jointly assessing both model capabilities and the validity of evaluation protocols themselves.

audio-language modelsautomatic judgingevaluation validity

Hot Scholars

HY

Hung-yi Lee

National Taiwan University
deep learningspoken language understandingspeech processing
SW

Shinji Watanabe

Carnegie Mellon University
Speech recognitionSpeech processingSpeech enhancementSpeech translation
YT

Yu Tsao

Research Fellow (Professor), Deputy Director, CITI, Academia Sinica
Assistive Oral Communication TechnologiesSpeech EnhancementVoice ConversionSpeech Assessment
HM

Hsin-Min Wang

Research Fellow/Professor, Institute of Information Sience, Academia Sinica
Spoken Language ProcessingNatural Language ProcessingMultimedia Information RetrievalMachine Learning
MB

Michael Beigl

Professor for Informatics, Karlsruhe Institute of Technology (KIT)
Ubiquitous ComputingWearable ComputingHealth & Activity Recognition using AIInternet of Things