Guiding Large Language Models to Post-Edit Machine Translation with Error Annotations

📅 2024-04-11
🏛️ NAACL-HLT
📈 Citations: 22
✨ Influential: 1
📄 PDF
🤖 AI Summary
Existing post-editing methods for machine translation fail to fully harness the capabilities of large language models (LLMs). This paper proposes a novel LLM-based post-editing framework that integrates fine-grained MQM error annotations with LLMs: for the first time, MQM quality labels are incorporated as interpretable external feedback into both prompting and supervised fine-tuning of LLaMA-2, enabling error-driven, precise editing. The method comprises MQM annotation parsing, multilingual modeling (Chinese–English, English–German, English–Russian), and feedback-aware instruction tuning. Experiments demonstrate consistent improvements over baselines across TER, BLEU, and COMET metrics; human evaluation confirms substantial gains in translation quality, while fine-tuning significantly enhances the model’s efficiency in leveraging fine-grained feedback. The core contribution is a principled, interpretable, and learnable MQM–LLM collaborative post-editing paradigm.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsHumans and AI: Human-in-the-loop Machine Learning

Application Category

Economics, Online Markets and Human Computation: LLM based quality controls for crowd workUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Machine Translation (MT) remains one of the last NLP tasks where large language models (LLMs) have not yet replaced dedicated supervised systems. This work exploits the complementary strengths of LLMs and supervised MT by guiding LLMs to automatically post-edit MT with external feedback on its quality, derived from Multidimensional Quality Metric (MQM) annotations. Working with LLaMA-2 models, we consider prompting strategies varying the nature of feedback provided and then fine-tune the LLM to improve its ability to exploit the provided guidance. Through experiments on Chinese-English, English-German, and English-Russian MQM data, we demonstrate that prompting LLMs to post-edit MT improves TER, BLEU and COMET scores, although the benefits of fine-grained feedback are not clear. Fine-tuning helps integrate fine-grained feedback more effectively and further improves translation quality based on both automatic and human evaluation.
Problem

Research questions and friction points this paper is trying to address.

Guiding LLMs to post-edit machine translation outputs
Integrating MQM error annotations as external feedback
Improving translation quality through fine-tuned error correction
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-guided post-editing with MQM annotations
Fine-tuning LLaMA-2 for error correction
Prompting strategies with quality feedback
🔎 Similar Papers
No similar papers found.