From Speech to Editable Concepts: Probing Emotion Recognition with Concept Bottleneck Models

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of interpretability in large language models (LLMs) for speech emotion recognition (SER), where aggregated metrics obscure the influence of individual concepts on specific predictions. To this end, this work introduces concept bottleneck models to the SER domain for the first time, integrating multimodal concepts—including transcriptions and acoustic descriptors—to systematically investigate LLM prediction dependencies and debiasing strategies. Our findings reveal transcription bias in zero-shot scenarios and demonstrate that fine-tuning effectively mitigates such biases while improving Macro-F1 scores. Furthermore, we show that removing specific acoustic concepts significantly alters individual predictions, thereby overcoming the limitations of conventional evaluation paradigms. Collectively, this research establishes a novel framework for advancing interpretability and fairness in SER.
📝 Abstract
Speech emotion recognition (SER) is the task of assigning emotion labels to utterances. Early systems relied on acoustic features, whereas recent approaches combine multiple modalities, most commonly speech and text. Still, performance remains poor on many datasets. Large language models (LLMs) have therefore attracted interest for SER, as they can process diverse inputs jointly with instructions. However, direct audio input raises questions of explainability. To address similar questions in image classification, concept bottleneck models were introduced. This work adapts concept bottlenecks to SER to examine how individual predictions depend on transcripts, acoustic descriptions and speaker attributes. Experiments test three LLMs on CREMA-D, IEMOCAP and MELD, with concepts extracted by separate tools. On scripted corpora, LLMs are strongly biased towards the transcript in the zero-shot setting, which lowers Macro-F1 from 27.8 to 5.8 on CREMA-D. Fine-tuning removes this bias, and the transcript raises Macro-F1 from 41.8 to 45.1. Removing speech rate changes 48% of Neutral predictions to Disgust on CREMA-D; removing intensity level on MELD changes predictions despite little change in Macro-F1. These findings show that aggregate performance changes alone do not capture the effects of concept removal on individual predictions.
Problem

Research questions and friction points this paper is trying to address.

Speech Emotion Recognition
Explainability
Large Language Models
Concept Bottleneck Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Concept Bottleneck Models
Speech Emotion Recognition
Large Language Models
Explainability
Concept Intervention
🔎 Similar Papers
No similar papers found.