MGhana-ST: A Low-Resource Speech Translation Dataset for Ghanaian Languages and an Analysis of Multilingual Training Trade-offs

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the scarcity of speech translation data for low-resource Ghanaian languages and the trade-offs of multilingual training under extreme data paucity. We introduce MGhana-ST, the first high-quality Ghanaian speech translation dataset comprising 16.1 hours of audio, constructed using the Whisper-small model with native speaker annotation. By comparing monolingual and multilingual joint training, we propose a seed-based statistical significance analysis method. Our experiments reveal that random seed variance substantially affects low-resource evaluation. Furthermore, flat multilingual training yields no performance gains and instead degrades results for Ewe and Fante, demonstrating that apparent positive transfer under conventional multilingual assumptions can be spurious due to noisy monolingual baselines. These findings challenge established assumptions regarding multilingual advantages in extremely low-resource scenarios. The dataset is publicly released.
📝 Abstract
We present MGhana-ST, a speech translation dataset for four low-resource Ghanaian language varieties: Ga, Twi (Akuapem and Asante), Ewe, and Fante. MGhana-ST is an ongoing annotation effort; the experiments here use a fixed subset of about 16.1 hours of paired speech and English translations. The audio is curated from two existing Ghanaian speech resources. Unlike in those resources, the English translations are produced directly from audio by 37 native-speaker annotators and include verbal and non-verbal event annotations. Using Whisper-small, we compare monolingual and multilingual training under severe data scarcity, reporting means over three seeds. Flat multilingual training benefits no variety in this regime. Ga and Twi are unchanged within seed variance (+0.51 and +0.06 BLEU against monolingual standard deviations of 1.63 and 2.20), while Ewe declines by 6.99 BLEU and Fante by 5.11. The degrading varieties are Ewe, which is linguistically distinct and drawn from a different source corpus, and Fante, the least-resourced. Comparing empirical cross-lingual transfer with typology-based similarity, we find that transfer BLEU identifies closely interacting language pairs better than URIEL similarity, though neither predicts which varieties benefit from joint training. We also report a methodological finding. An earlier single-run analysis found positive transfer for three of four varieties; this did not survive replication across seeds. For Ga and Twi, monolingual baselines trained on 1.6 to 6.2 hours of audio have seed standard deviations roughly five and thirty times those of the multilingual models (0.35 and 0.07 BLEU). When the monolingual condition is noisier, a single-run comparison can show apparent transfer of this size from seed variation alone. We release MGhana-ST to support research on African language speech technology and low-resource speech translation.
Problem

Research questions and friction points this paper is trying to address.

low-resource speech translation
Ghanaian languages
multilingual training trade-offs
cross-lingual transfer
evaluation reproducibility
Innovation

Methods, ideas, or system contributions that make the work stand out.

Low-Resource Speech Translation
Multilingual Training Trade-offs
Cross-lingual Transfer
African Languages
Reproducibility
🔎 Similar Papers
No similar papers found.
F
Frank Lawrence Nii Adoquaye Acquaye
Ashesi University, Berekuso, Ghana
E
Eric George Parakal
HSE University, Moscow, Russia
J
Jesse Johnson
AdwumaTech AI, Accra, Ghana
K
Kishankumar Bhimani
GIFT International Fintech Institute, Gandhinagar, India
J
Jochebed Afua Basil
Ashesi University, Berekuso, Ghana