🤖 AI Summary
This study addresses the vulnerability of audio classifiers to learning spurious correlations, or shortcuts, which compromises their generalization in real-world scenarios. To mitigate this issue, the authors propose an interpretability auditing framework that, for the first time, integrates temporal explanations with large language models. Specifically, the method isolates critical explanatory segments, leverages large audio-language models to generate textual descriptions, and automatically extracts recurring semantic concepts to identify potential shortcut behaviors. The proposed approach successfully reveals spurious associations, such as relying on "laughter" to predict "applause," thereby validating its reliability and effectiveness in uncovering shortcut learning within audio classification models.
📝 Abstract
Correlations between events in machine learning datasets may result in shortcut learning, where models learn to predict the target event based on the presence of a correlated event. When these correlations are spurious -- arising from data collection artifacts -- models are likely to perform poorly in practice. We propose a pipeline to uncover shortcut learning in audio classifiers by discovering recurring concepts in their temporal explanations. Specifically, we isolate audio segments that explain classifier decisions, caption them with an ensemble of Large Audio-Language Models, and use a Large Language Model to extract recurring concepts. The resulting concepts can be audited by humans to uncover potential shortcut learning. We evaluate our framework using datasets curated from AudioSet Strong, controlling for the presence or absence of spurious correlations. Results show that this approach reliably uncovers learned shortcuts, such as the model relying on the presence of"laughter"to predict"applause".