Machine Translation Evaluation: A Survey

📅 2016-05-15
🏛️ arXiv.org
📈 Citations: 13
✨ Influential: 1
📄 PDF
🤖 AI Summary
To address the lack of systematic meta-evaluation in machine translation (MT) assessment, this paper proposes the first unified taxonomy that integrates human evaluation (readability, adequacy, post-editing effort) with automatic metrics (based on morphological, syntactic, and semantic features), while incorporating emerging paradigms such as quality estimation (QE). We introduce a fine-grained classification framework for syntactic and semantic features—including dependency parsing, semantic role labeling, textual entailment, and pre-trained language models—and design a customizable automatic metric framework tailored for large language models. Additionally, we formalize sample-size requirements for human evaluation. The resulting taxonomy constitutes the most comprehensive, structured, and up-to-date MT evaluation framework to date, significantly enhancing methodological rigor and experimental reproducibility. It provides both theoretical foundations and practical guidelines for robust MT evaluation and QE research.
📝 Abstract
This paper introduces the state-of-the-art machine translation (MT) evaluation survey that contains both manual and automatic evaluation methods. The traditional human evaluation criteria mainly include the intelligibility, fidelity, fluency, adequacy, comprehension, and informativeness. The advanced human assessments include task-oriented measures, post-editing, segment ranking, and extended criteriea, etc. We classify the automatic evaluation methods into two categories, including lexical similarity scenario and linguistic features application. The lexical similarity methods contain edit distance, precision, recall, F-measure, and word order. The linguistic features can be divided into syntactic features and semantic features respectively. The syntactic features include part of speech tag, phrase types and sentence structures, and the semantic features include named entity, synonyms, textual entailment, paraphrase, semantic roles, and language models. Subsequently, we also introduce the evaluation methods for MT evaluation including different correlation scores, and the recent quality estimation (QE) tasks for MT. This paper differs from the existing works cite{GALEprogram2009,EuroMatrixProject2007} from several aspects, by introducing some recent development of MT evaluation measures, the different classifications from manual to automatic evaluation measures, the introduction of recent QE tasks of MT, and the concise construction of the content.
Problem

Research questions and friction points this paper is trying to address.

Evaluating Machine Translation (MT) quality metrics evolution
Assessing human parity gaps in Neural Machine Translation (NMT)
Meta-analysis of automatic and human evaluation methods for MT
Innovation

Methods, ideas, or system contributions that make the work stand out.

Neural models leverage large parallel corpora
Pre-trained language models customize evaluation metrics
Statistical confidence estimates human evaluation samples
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
The University of Manchester | Leiden University | Logrus Global LLC