Extending Multilingual Machine Translation through Imitation Learning

📅 2023-11-14
🏛️ arXiv.org
📈 Citations: 5
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the challenge of low-resource language integration, this paper proposes a method for extending large-scale multilingual neural machine translation (MNMT) models using only bilingual corpora between a new language and English. The approach enables bidirectional translation between the new language and all pre-existing languages while mitigating catastrophic forgetting. Its core innovation is the first application of imitation learning to MNMT: leveraging English as a pivot to construct pseudo-multilingual parallel data and distilling the original model’s output distribution across the multilingual target space. This strategy effectively suppresses translation copying and off-target errors. Experiments demonstrate that the method achieves an average BLEU gain of over 3.0 for the new language, with degradation in existing language performance constrained to less than 0.5 BLEU—significantly outperforming baseline approaches.
📝 Abstract
Despite the growing variety of languages supported by existing multilingual neural machine translation (MNMT) models, most of the world's languages are still being left behind. We aim to extend large-scale MNMT models to a new language, allowing for translation between the newly added and all of the already supported languages in a challenging scenario: using only a parallel corpus between the new language and English. Previous approaches, such as continued training on parallel data including the new language, suffer from catastrophic forgetting (i.e., performance on other languages is reduced). Our novel approach Imit-MNMT treats the task as an imitation learning process, which mimicks the behavior of an expert, a technique widely used in the computer vision area, but not well explored in NLP. More specifically, we construct a pseudo multi-parallel corpus of the new and the original languages by pivoting through English, and imitate the output distribution of the original MNMT model. Extensive experiments show that our approach significantly improves the translation performance between the new and the original languages, without severe catastrophic forgetting. We also demonstrate that our approach is capable of solving copy and off-target problems, which are two common issues existence in current large-scale MNMT models.
Problem

Research questions and friction points this paper is trying to address.

Extend multilingual translation to new languages using limited data
Mitigate catastrophic forgetting when incorporating new language pairs
Apply imitation learning to generate pseudo-parallel corpora for training
Innovation

Methods, ideas, or system contributions that make the work stand out.

Imit-MNMT uses imitation learning for language extension
Generates pseudo-parallel corpora via expert model
Employs data distribution and behavior imitation strategies
🔎 Similar Papers
No similar papers found.