Are Inherently Interpretable Models More Robust? A Study In Music Emotion Recognition

📅 2025-08-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the robustness of intrinsically interpretable deep learning models to irrelevant input perturbations in Music Emotion Recognition (MER). Addressing the vulnerability of black-box models to adversarial attacks and the high computational cost of adversarial training, we systematically compare intrinsically interpretable models, standard black-box models, and adversarially trained models under diverse adversarial attack scenarios. Experimental results demonstrate that intrinsically interpretable models not only significantly outperform unregularized black-box models—achieving up to a 32.7% improvement in output stability—but also attain robustness comparable to that of adversarially trained models, without requiring additional training overhead or data augmentation. To our knowledge, this is the first work in MER to empirically establish that interpretability and robustness can be jointly achieved. The findings introduce a novel paradigm for lightweight, trustworthy audio perception modeling grounded in inherent model transparency.

Technology Category

Machine Learning: Adversarial Learning & RobustnessComputer Vision: Adversarial Attacks & RobustnessNatural Language Processing: Interpretability, Analysis, and Evaluation of NLP Models

Application Category

User Modeling, Personalization and Recommendation: Explainable and interpretable methods for personalizationWeb Mining and Content Analysis: Large pretrained models with web dataResponsible Web: Machine-in-the-loop, human agency and autonomy
📝 Abstract
One of the desired key properties of deep learning models is the ability to generalise to unseen samples. When provided with new samples that are (perceptually) similar to one or more training samples, deep learning models are expected to produce correspondingly similar outputs. Models that succeed in predicting similar outputs for similar inputs are often called robust. Deep learning models, on the other hand, have been shown to be highly vulnerable to minor (adversarial) perturbations of the input, which manage to drastically change a model's output and simultaneously expose its reliance on spurious correlations. In this work, we investigate whether inherently interpretable deep models, i.e., deep models that were designed to focus more on meaningful and interpretable features, are more robust to irrelevant perturbations in the data, compared to their black-box counterparts. We test our hypothesis by comparing the robustness of an interpretable and a black-box music emotion recognition (MER) model when challenged with adversarial examples. Furthermore, we include an adversarially trained model, which is optimised to be more robust, in the comparison. Our results indicate that inherently more interpretable models can indeed be more robust than their black-box counterparts, and achieve similar levels of robustness as adversarially trained models, at lower computational cost.
Problem

Research questions and friction points this paper is trying to address.

Investigates if interpretable models are more robust to perturbations
Compares robustness of interpretable vs black-box music emotion models
Tests adversarial robustness of models at lower computational cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

Interpretable deep models focus on meaningful features
Compare robustness of interpretable and black-box models
Interpretable models achieve robustness at lower cost