🤖 AI Summary
This study addresses the limitation of existing spoken language models in supporting global applications for low-resource languages due to a lack of multilingual instruction-tuning data. We propose a data construction strategy that combines large-scale synthesis with human verification, and release MULTISPEECHQA, a ten-million-scale multilingual question-answering dataset covering 23 languages, alongside the evaluation benchmark MULTISPEECH-BENCH. Through instruction tuning and multi-task evaluation based on Qwen 2.5-Omni, our results demonstrate that high-quality synthetic data constitutes an effective, low-cost pathway for enhancing the multilingual capabilities of such models. The fine-tuned model achieves substantial performance improvements, with its cascaded system outperforming mainstream open-source models. These findings validate the innovative value of the proposed approach for multilingual speech understanding.
📝 Abstract
Speech Language Models (SLMs) that understand spoken language questions support only a few high-resource languages, limiting access to millions of people worldwide. This gap stems from the scarcity of multilingual speech instruction-tuning datasets. We present MULTISPEECHQA, a large-scale, synthetically generated and human-verified dataset comprising 9200 hours of 10.8 million spoken question-answer pairs in 23 typologically diverse languages. Using MULTISPEECHQA, we also introduce MULTISPEECH-BENCH, a multi-task benchmark for evaluating SLM performance on 23 languages. We compare the performance of a cascading system to open-weight and closed SLMs on MULTISPEECH-BENCH and find that the cascading system outperforms open-weight SLMs but not all closed SLMs. We use MULTISPEECHQA to finetune Qwen 2.5-Omni, which improves its performance on our benchmark. Our findings show that high-quality synthetic datasets offer a cheap solution to improving the multilingual capabilities of SLMs.