🤖 AI Summary
This study addresses spoken Romanian dialect identification, a challenging task due to regional phonetic and prosodic variation. We introduce MoRoVoc, the largest publicly available Romanian dialect speech dataset to date (93 hours, 88,192 utterances), covering major dialect regions across Romania and Moldova. To disentangle dialect-specific features from speaker demographics—particularly gender and age—we propose a novel multi-objective adversarial training framework integrating meta-learning with dynamic coefficient adjustment. This is the first work to apply meta-learning to demographic attribute disentanglement in speech dialect recognition. Built upon the Wav2Vec2 architecture, our method achieves 78.21% accuracy on dialect classification and improves gender classification to 93.08% using Wav2Vec2-Large with dual adversarial objectives. The approach significantly enhances model robustness and cross-domain generalization, demonstrating strong performance in mitigating demographic bias while preserving dialect discriminability.
📝 Abstract
This paper introduces MoRoVoc, the largest dataset for analyzing the regional variation of spoken Romanian. It has more than 93 hours of audio and 88,192 audio samples, balanced between the Romanian language spoken in Romania and the Republic of Moldova. We further propose a multi-target adversarial training framework for speech models that incorporates demographic attributes (i.e., age and gender of the speakers) as adversarial targets, making models discriminative for primary tasks while remaining invariant to secondary attributes. The adversarial coefficients are dynamically adjusted via meta-learning to optimize performance. Our approach yields notable gains: Wav2Vec2-Base achieves 78.21% accuracy for the variation identification of spoken Romanian using gender as an adversarial target, while Wav2Vec2-Large reaches 93.08% accuracy for gender classification when employing both dialect and age as adversarial objectives.