Score
Design and run controlled subjective evaluation protocols that present audio stimuli to screened human listeners and collect perceptual responses such as ratings, preferences, or categorical judgments. Build the test materials and procedures, recruit and screen participants, manage stimulus presentation and data capture, and analyze listener data statistically to compare systems, estimate perceptual attributes, and validate audio-processing methods.
In text-to-audio (TTA) generation, evaluating text–audio relevance has long relied on costly human assessments or objective metrics of questionable validity (e.g., CLAPScore). To address this, we introduce RELATE—the first open-source, human-annotated subjective evaluation dataset for TTA relevance assessment, covering diverse acoustic categories and providing fine-grained relevance scores. Leveraging RELATE, we train an end-to-end deep learning model to predict human relevance judgments automatically. Experiments demonstrate that our model significantly outperforms CLAPScore across all sound categories, achieves consistently high performance, and exhibits strong agreement with human raters (Spearman ρ > 0.72). This work establishes the first standardized subjective benchmark for TTA relevance evaluation and provides a reliable, automated assessment tool—thereby filling two critical gaps in the field.
This study addresses a critical yet overlooked issue in evaluating large audio language models: their potential reliance on cues from evaluation protocols—such as provided labels or reference information—rather than genuine comprehension of audio content, which can artificially inflate human-model agreement. To probe such protocol-level shortcuts, the authors introduce a novel auditing framework employing three adversarial manipulations: feature blueprint substitution, reference information perturbation, and option order swapping. Applying this approach across six prominent models and four speech attributes, the experiments reveal substantial shortcut dependence; for instance, emotion recognition accuracy drops below 0.10 for several models, and Qwen3-Omni-Thinking consistently selects the same position in A/B tests regardless of content. These findings underscore the necessity of jointly assessing both model capabilities and the validity of evaluation protocols themselves.
This study investigates how listeners perceive AI-generated music, focusing on the effects of composer identity labels, personality traits, musical preferences, and perceived humanness on acceptance and affective responses. Employing a mixed-methods design, it utilized multi-genre AI-composed music as stimuli, combined with standardized psychometric scales, real-time affective assessment, and in-depth thematic analysis. Results indicate that general attitudes toward AI constitute the strongest predictor of musical liking and emotional intensity; qualitative findings further identify ethical concerns, cultural context, and usage scenarios as three core evaluative dimensions shaping listener judgments. Moving beyond a technocentric paradigm, this research is the first to systematically integrate individual differences with sociocultural factors, thereby establishing an empirical foundation and methodological framework for human-AI musical interaction. It advances scholarship on AI art acceptance from the question “Can AI music be accepted?” to “Why is it accepted—or not—in specific ways?”
Existing audiovisual quality assessment datasets are limited in scale, lack diversity in content and quality degradation types, and provide only holistic scores, thereby hindering research on multimodal perception mechanisms. To address these limitations, this work proposes a crowdsourced subjective evaluation framework that transcends traditional laboratory constraints, integrating a systematic data sampling strategy with a multidimensional annotation scheme. This approach yields YT-NTU-AVQ, the largest and most diverse audiovisual quality assessment dataset to date, comprising 1,620 user-generated videos spanning a broad spectrum of semantic scenarios and quality levels. The dataset and associated platform code have been publicly released, significantly advancing the study and development of multimodal perceptual modeling.
In real-world acoustic scenes, users struggle to manipulate unseparated mixed sound sources. To address this, we propose the first end-to-end, text-driven framework for real-time sound field editing, enabling direct, joint manipulation of multiple concurrent sources—such as “reduce air-conditioner noise and enhance speech”—via natural language instructions, without explicit source separation. Our method integrates large language model–based semantic parsing with a differentiable spectrogram decomposition–filtering–reconstruction architecture, supporting open-vocabulary, zero-shot editing. It is trained on a newly curated dataset comprising 160 hours of audio and 100,000 audio–text pairs. Experiments demonstrate significant improvements in source extraction, suppression, and level control: +3.2 dB in SI-SNR and +0.11 in STOI, while maintaining strong robustness and generalization across complex mixtures with 2–5 overlapping sources.
This study addresses the limitation of objective audio quality assessment methods that yield only point estimates without uncertainty quantification. We propose a confidence interval prediction approach integrating a PEAQ frontend with bagging ensemble regression. By calibrating against subjective data to generate score distributions, the method achieves performance comparable to end-to-end models using lightweight features and public datasets. Experimental results demonstrate that the proposed approach maintains mean opinion score (MOS) prediction accuracy while significantly improving alignment with subjective score distributions and confidence interval coverage rates. Furthermore, it supports approximate pairwise comparisons. Overall, this work provides a reliable, uncertainty-aware framework for audio quality evaluation.
研究通过对比试验和对话编码,发现职业反思代理在执行系统提示时存在不一致,尤其是对决策的要求增加了参与者的疑虑。
This study addresses the lag in human evaluation tooling for generative models and the absence of open-source, self-hosted platforms supporting controlled experiments by developing a multimodal web-based evaluation system. The platform employs a browser-side editing and distribution architecture, incorporating built-in significance testing and Bradley-Terry scoring algorithms. Furthermore, it integrates GDPR-compliant data auditing and access control modules to support complex experimental logic and machine-readable specification export. By releasing the complete codebase as open source, this work provides standardized infrastructure that enables researchers to efficiently conduct privacy-compliant, controlled human evaluation experiments.
This work addresses a critical gap in the evaluation of audio language models (ALMs) as speech evaluators: their purported reliance on paralinguistic cues—such as emotion and prosody—may be illusory. To investigate, the authors propose a counterfactual auditing framework that constructs contrastive audio samples preserving textual content while varying only paralinguistic features. By jointly analyzing native judgments and performance on a recoverability control task—and further disentangling perceptual encoding from response mapping—the framework localizes the sources of model failure. Systematic evaluation across Gemini, GPT, and open-source ALMs reveals that high contrastive accuracy often masks unreliable native judgments and that heterogeneous failure modes can coexist under similar aggregate performance. These findings underscore the necessity of fine-grained behavioral auditing beyond conventional accuracy metrics and establish a new paradigm for trustworthy evaluation of audio language models.
This study addresses the challenge of establishing reliable ground truth in audio scene description, where annotator subjectivity introduces significant ambiguity. To resolve this, we propose defining sound salience through speakers' audible reactions. Methodologically, we construct CARES, a synthetic benchmark corpus comprising ten thousand dual-speaker scenarios generated via controlled scene design and large language model-driven dialogue synthesis, thereby effectively eliminating annotation ambiguity. We evaluate six audio language models on this benchmark. Our experiments reveal that while existing models can recognize environmental sounds, they struggle to accurately capture speakers' behavioral responses to these sounds. This finding exposes a critical bottleneck in the field, highlighting the substantial gap between basic acoustic recognition and modeling human auditory interaction within complex audio scenes.