AccentCL: Robust Accent Classification with Incremental Expansion

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of incrementally expanding accent classification systems hindered by fixed label sets, class imbalance, and cross-domain shifts. We propose AccentCL, a framework that employs a frozen Whisper-Large-v3 feature encoder combined with replay-based continual learning to seamlessly accommodate novel accent classes. Furthermore, we introduce an imbalance-aware loss, domain mean alignment, and an old-to-new margin loss to effectively suppress over-prediction of new classes while enhancing cross-domain robustness. Experimental results demonstrate that the proposed model achieves a balanced accuracy of 77.1% on a five-class task. Notably, when Spanish and Chinese accents are incrementally introduced, performance on base classes remains undegraded, and the F1 score for novel classes improves to 83.3%.
📝 Abstract
Accent classifiers are typically trained with a fixed label inventory and cannot accommodate new accent categories as new data becomes available. Moreover, accented speech corpora often exhibit substantial class imbalance and/or domain shift due to differences in recording conditions across corpora. We present AccentCL, a class-incremental learning framework for English accent classification that is robust to class imbalance and cross-corpus domain shift. AccentCL extracts multi-layer representations from a frozen Whisper-Large-v3 encoder, optimized with an imbalance-aware cross-entropy loss to reduce bias toward the majority accent classes and a domain mean alignment loss that minimizes distributional mean shift across training corpora. The label space is then expanded via replay-based continual learning, using the frozen base model for knowledge retention and an old-to-new margin loss to reduce overprediction on newly added classes. On a five-class accent classification task, AccentCL achieves 77.1% balanced accuracy and a 76.9% macro-averaged F1 score. We further evaluate the model's ability to incrementally incorporate two new accent categories: Spanish-accented and Chinese-accented English. When adding Spanish-accented English to the pretrained model, AccentCL attains an F1 of 83.3% on the new class while retaining 77.3% balanced accuracy on the base classes. When subsequently adding Chinese-accented English, it achieves 61.8% F1 on the new class while preserving 77.6% balanced accuracy on the previously learned classes. These results show that AccentCL enables robust regional accent classification while allowing new accent categories to be added without full retraining.
Problem

Research questions and friction points this paper is trying to address.

accent classification
class-incremental learning
class imbalance
domain shift
Innovation

Methods, ideas, or system contributions that make the work stand out.

Class-Incremental Learning
Accent Classification
Domain Shift
Class Imbalance
Continual Learning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Mu-Ruei Tseng
Department of Computer Science and Engineering, Texas A&M University
W
Waris Quamer
Department of Computer Science and Engineering, Texas A&M University
G
Ghady Nasrallah
Department of Computer Science and Engineering, Texas A&M University
Ricardo Gutierrez-Osuna
Ricardo Gutierrez-Osuna
Texas A&M University, Computer Science and Engineering
Speech generationdigital healthwearable sensorsmachine learningchemometrics