Institution profile

Ashesi University

Academic institutionafrica · gh
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

MGhana-ST: A Low-Resource Speech Translation Dataset for Ghanaian Languages and an Analysis of Multilingual Training Trade-offs

Sep 30, 2026

This study addresses the scarcity of speech translation data for low-resource Ghanaian languages and the trade-offs of multilingual training under extreme data paucity. We introduce MGhana-ST, the first high-quality Ghanaian speech translation dataset comprising 16.1 hours of audio, constructed using the Whisper-small model with native speaker annotation. By comparing monolingual and multilingual joint training, we propose a seed-based statistical significance analysis method. Our experiments reveal that random seed variance substantially affects low-resource evaluation. Furthermore, flat multilingual training yields no performance gains and instead degrades results for Ewe and Fante, demonstrating that apparent positive transfer under conventional multilingual assumptions can be spurious due to noisy monolingual baselines. These findings challenge established assumptions regarding multilingual advantages in extremely low-resource scenarios. The dataset is publicly released.

0 citationsRead paper

VAMAE: Vessel-Aware Masked Autoencoders for OCT Angiography

Apr 07, 2026

This work addresses the challenge of self-supervised representation learning in OCTA images, where sparse vasculature and strong topological constraints hinder effective feature learning. To this end, the authors propose a vessel-aware masked autoencoder framework that integrates vessel saliency with skeleton priors to devise an anatomy-guided, non-uniform masking strategy. By jointly optimizing multi-objective reconstruction tasks, the method simultaneously preserves vascular appearance, structural continuity, and topological fidelity, thereby enabling geometry-aware learning of vessel connectivity and branching patterns. Experiments on the OCTA-500 benchmark demonstrate that the proposed approach significantly outperforms standard masked autoencoders, with particularly notable gains in label-scarce settings.

0 citationsRead paper

AFRICAPTION: Establishing a New Paradigm for Image Captioning in African Languages

Oct 20, 2025

African languages—representing a significant portion of the world’s linguistic diversity—are severely underrepresented in multimodal AI, particularly in image captioning, due to scarce annotated data and limited model support. Method: This work introduces the first large-scale vision-to-language framework for 20 African languages. It constructs a semantically aligned, high-quality multilingual image-caption dataset; designs a dynamic quality assurance pipeline integrating context-aware translation, model ensembling (SigLIP + NLLB-200), and adaptive token replacement; and develops a unified, 0.5B-parameter vision-to-text architecture optimized for low-resource settings. Contribution/Results: We release the first open-source, African-language–focused image captioning dataset and corresponding pre-trained models. Our framework establishes a new multilingual generation paradigm that balances accuracy and scalability, achieving substantial performance gains on cross-modal tasks for low-resource languages. This advances inclusive, equitable multimodal AI development and sets a foundation for future research in under-resourced language modalities.

0 citationsRead paper
Recent publications

Latest Papers

MGhana-ST: A Low-Resource Speech Translation Dataset for Ghanaian Languages and an Analysis of Multilingual Training Trade-offs

Sep 30, 2026

This study addresses the scarcity of speech translation data for low-resource Ghanaian languages and the trade-offs of multilingual training under extreme data paucity. We introduce MGhana-ST, the first high-quality Ghanaian speech translation dataset comprising 16.1 hours of audio, constructed using the Whisper-small model with native speaker annotation. By comparing monolingual and multilingual joint training, we propose a seed-based statistical significance analysis method. Our experiments reveal that random seed variance substantially affects low-resource evaluation. Furthermore, flat multilingual training yields no performance gains and instead degrades results for Ewe and Fante, demonstrating that apparent positive transfer under conventional multilingual assumptions can be spurious due to noisy monolingual baselines. These findings challenge established assumptions regarding multilingual advantages in extremely low-resource scenarios. The dataset is publicly released.

0 citationsRead paper

VAMAE: Vessel-Aware Masked Autoencoders for OCT Angiography

Apr 07, 2026

This work addresses the challenge of self-supervised representation learning in OCTA images, where sparse vasculature and strong topological constraints hinder effective feature learning. To this end, the authors propose a vessel-aware masked autoencoder framework that integrates vessel saliency with skeleton priors to devise an anatomy-guided, non-uniform masking strategy. By jointly optimizing multi-objective reconstruction tasks, the method simultaneously preserves vascular appearance, structural continuity, and topological fidelity, thereby enabling geometry-aware learning of vessel connectivity and branching patterns. Experiments on the OCTA-500 benchmark demonstrate that the proposed approach significantly outperforms standard masked autoencoders, with particularly notable gains in label-scarce settings.

0 citationsRead paper

AFRICAPTION: Establishing a New Paradigm for Image Captioning in African Languages

Oct 20, 2025

African languages—representing a significant portion of the world’s linguistic diversity—are severely underrepresented in multimodal AI, particularly in image captioning, due to scarce annotated data and limited model support. Method: This work introduces the first large-scale vision-to-language framework for 20 African languages. It constructs a semantically aligned, high-quality multilingual image-caption dataset; designs a dynamic quality assurance pipeline integrating context-aware translation, model ensembling (SigLIP + NLLB-200), and adaptive token replacement; and develops a unified, 0.5B-parameter vision-to-text architecture optimized for low-resource settings. Contribution/Results: We release the first open-source, African-language–focused image captioning dataset and corresponding pre-trained models. Our framework establishes a new multilingual generation paradigm that balances accuracy and scalability, achieving substantial performance gains on cross-modal tasks for low-resource languages. This advances inclusive, equitable multimodal AI development and sets a foundation for future research in under-resourced language modalities.

0 citationsRead paper