MGhana-ST: A Low-Resource Speech Translation Dataset for Ghanaian Languages and an Analysis of Multilingual Training Trade-offs
This study addresses the scarcity of speech translation data for low-resource Ghanaian languages and the trade-offs of multilingual training under extreme data paucity. We introduce MGhana-ST, the first high-quality Ghanaian speech translation dataset comprising 16.1 hours of audio, constructed using the Whisper-small model with native speaker annotation. By comparing monolingual and multilingual joint training, we propose a seed-based statistical significance analysis method. Our experiments reveal that random seed variance substantially affects low-resource evaluation. Furthermore, flat multilingual training yields no performance gains and instead degrades results for Ewe and Fante, demonstrating that apparent positive transfer under conventional multilingual assumptions can be spurious due to noisy monolingual baselines. These findings challenge established assumptions regarding multilingual advantages in extremely low-resource scenarios. The dataset is publicly released.