🤖 AI Summary
This work addresses the challenge of noisy labels and poor domain alignment in large-scale weakly supervised speech recognition data, which often limits model performance. The authors propose a three-stage training strategy: first, pretraining on the full weakly supervised dataset; second, refining the model through a second pretraining phase using a high-quality subset selected based on character error rate (CER); and third, fine-tuning on samples from this subset that are acoustically similar to the target domain, as measured by an acoustic similarity metric. By jointly incorporating data filtering and domain-relevant sample selection without introducing external data, the method significantly improves recognition accuracy. Experiments on 90,000 hours of Japanese weakly supervised data demonstrate CER reductions of up to 6.4% and 4.0% under the proposed approach.
📝 Abstract
Leveraging large-scale weakly supervised datasets is crucial to train robust end-to-end automatic speech recognition (ASR) models. However, such datasets often contain noisy labels and lack domain specificity, limiting their effectiveness. To address these issues and make better use of weakly supervised datasets, we propose a novel training approach incorporating data filtering and selection. Our approach consists of three steps: pretraining on the entire dataset, continued pretraining on a filtered subset based on character error rate (CER), and fine-tuning on a small number of acoustically similar samples to the target domain, selected from the filtered subset. In experiments with a 90,000-hour weakly supervised Japanese dataset, the proposed filtering and selection methods synergistically reduced CER by up to 6.4% and 4.0%, respectively, even though these steps reused training samples already used in the first pretraining step.