A Study of ASR Adaptation and Representation Dimensionality Reduction in Persian Speech Emotion Recognition Using Whisper

📅 2026-05-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the scarcity of labeled data in speech emotion recognition for low-resource languages such as Persian by proposing an efficient framework that leverages frame-level embeddings extracted from the Whisper encoder. To substantially reduce computational overhead and model parameters, the approach employs parameter-free PCA for dimensionality reduction, followed by attention-based pooling and a lightweight classification head for emotion prediction. The study systematically evaluates the transferability of Persian automatic speech recognition (ASR) fine-tuning to emotion recognition, achieving superior performance on the ShEMO dataset while significantly lowering training latency and memory consumption. Experimental results indicate that the performance gains from ASR fine-tuning are limited, revealing inherent constraints in cross-lingual adaptation and the transfer of emotion-specific representations.
📝 Abstract
Speech Emotion Recognition (SER) in low-resource languages remains a challenging problem due to limited labeled data. In this work, we study the use of Whisper for Persian SER with a particular focus on representation dimensionality reduction and language-specific model adaptation. We propose a SER framework in which frame-level embeddings extracted from the Whisper encoder are reduced in dimensionality using PCA, eliminating the need for learned projection layers and substantially reducing the number of trainable parameters. The reduced representations are aggregated using an attention-based pooling mechanism and classified with a lightweight prediction head. In addition, we investigate whether fine-tuning Whisper on a Persian automatic speech recognition (ASR) task improves downstream SER performance. Experiments conducted on the ShEMO dataset under a speaker-independent evaluation protocol show that PCA-based dimensionality reduction consistently improves emotion recognition performance while reducing training latency and memory usage. ASR fine-tuning yields only modest gains for SER, suggesting limited transfer from language adaptation to emotion-related representations under the evaluated conditions. These findings provide practical insights into the efficient use of large pretrained speech models for emotion recognition in low-resource languages.
Problem

Research questions and friction points this paper is trying to address.

Speech Emotion Recognition
low-resource languages
Persian
limited labeled data
Innovation

Methods, ideas, or system contributions that make the work stand out.

dimensionality reduction
Whisper
speech emotion recognition
low-resource languages
PCA
🔎 Similar Papers
A
Ali Shendabadi
Faculty of Intelligent Systems Engineering, College of Interdisciplinary Sciences and Technologies, University of Tehran, Tehran, Iran
P
Parnia Izadirad
Faculty of Intelligent Systems Engineering, College of Interdisciplinary Sciences and Technologies, University of Tehran, Tehran, Iran
Mostafa Salehi
Mostafa Salehi
Associate Professor, University of Tehran
Social Network and Media AnalysisNetwork Science