Low resource cross-modal alignment using HGNN to enhance speech representation

📅 2026-09-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种基于异构图神经网络和链接预测的低资源跨模态对齐方法,用于增强语音表示,减少了对大量训练数据的需求。
📝 Abstract
Speech-text space alignment is a multimodal representation learning method consisting to map different speech and text into a shared representation space, leading to enrichment of the representation of each modality. Proposed architectures, such as SAMU-XLSR, typically follow a student/teacher framework, with the goal of fine-tuning an audio encoder to produce representations that closely match those of the text. In this way a speech representation is semantically enriched. However, such systems generally require large amounts of training data and considerable computational resource, making them difficult to apply to low resources languages under frugal constraints. The present work proposes a data-efficient space alignment method based on Heterogeneous Graph Neural Networks and link prediction. The core idea is to leverage message passing to explicitly transfer information from the text modality to the speech modality, thereby reducing the need for large training datasets and intrinsically enriching the acoustic representations, all in a more interpretable manner. Although thoroughly explored for high-resource languages, word-level tasks in speech remain relevant for certain low-resource languages. Therefore, we conducted experiments on speech-text alignment at the word level using the TIMIT (English) dataset and Yemba (a Cameroonian language). Our approach yields results comparable to those of SAMU-XLSR, a state-of-the-art method, and even surpasses it for the Yemba language in the task of word retrieval, while using far fewer resources, demonstrating its power, frugality, and efficiency.
Problem

Research questions and friction points this paper is trying to address.

Low resource
Cross-modal alignment
Speech representation
Heterogeneous Graph Neural Networks
Data-efficient
Innovation

Methods, ideas, or system contributions that make the work stand out.

Heterogeneous Graph Neural Networks
link prediction
low resource languages
cross-modal alignment
speech representation
💼 Related Jobs
No related jobs found.
Y
Yannick Yomie Nzeuhang
Department of computer sciences, University of Yaounde I, Street, Yaounde, 812, Cameroon
M
Marie Tahon
IRD, UMMISCO, Street, Bondy, F-93143, France
P
Paulin Melatagia Yonta
LIUM, Le Mans Universit´e, Avenue Olivier Messiaen, Le Mans, 72085, France