🤖 AI Summary
This study addresses the issue of multimodal machine translation models neglecting visual information and exhibiting low visual sensitivity. To overcome this limitation, we propose an optimization method based on metric loss weighting. Specifically, by introducing Congruency-based Pointwise Cross-Mutual Information (Congruency-based PCXMI), our approach accurately identifies tokens that rely on visual context and dynamically increases their training loss weights, thereby enhancing the visual grounding capabilities of pretrained multimodal large language models. Experimental results demonstrate that the proposed method improves accuracy on the CoMMuTE dataset by more than seven percentage points compared to standard fine-tuning, while maintaining excellent general translation performance.
📝 Abstract
Multimodal Machine Translation aims to incorporate additional signal from non-textual modalities to improve translations by resolving ambiguities. While models, through multimodal fusion, are able to accept images related to the source text, they can ignore this information. Therefore, increasing their visual sensitivity remains an active research area. In this work, we introduce a training method, Metric-based Loss Weighting, that improves visual grounding of translations by increasing the loss function for tokens that benefit from the accompanying image. We identify these tokens using the Point-wise Cross-mutual Information (PCXMI) metric, which compares the model's output probabilities with and without visual context. We introduce a Congruency-based PCXMI metric and experimentally show that both metrics working in combination yield the best results. We evaluate our method by fine-tuning three pretrained Multimodal Large Language Models on the task of Image-guided Machine Translation for three language directions. Metric-based Loss Weighting outperforms other tested methods on the CoMMuTE contrastive dataset, improving accuracy by up to more than 7 percentage points compared to standard fine-tuning, while maintaining strong general translation performance.