Score
Designs and applies evaluation frameworks, automatic metrics, and system-comparison experiments to measure and compare machine translation outputs on dimensions such as fluency, adequacy, terminological accuracy, and error severity. Builds or configures translation systems to produce test outputs, creates and validates annotation schemes (e.g., for metaphor translation or error-severity), and conducts quantitative and qualitative analyses to quantify system performance, changes to linguistic phenomena, and perceived quality or post-editing effort.
To address the lack of systematic meta-evaluation in machine translation (MT) assessment, this paper proposes the first unified taxonomy that integrates human evaluation (readability, adequacy, post-editing effort) with automatic metrics (based on morphological, syntactic, and semantic features), while incorporating emerging paradigms such as quality estimation (QE). We introduce a fine-grained classification framework for syntactic and semantic features—including dependency parsing, semantic role labeling, textual entailment, and pre-trained language models—and design a customizable automatic metric framework tailored for large language models. Additionally, we formalize sample-size requirements for human evaluation. The resulting taxonomy constitutes the most comprehensive, structured, and up-to-date MT evaluation framework to date, significantly enhancing methodological rigor and experimental reproducibility. It provides both theoretical foundations and practical guidelines for robust MT evaluation and QE research.
Professional-domain machine translation often neglects communicative goals and client requirements—i.e., translation norms—leading to outputs misaligned with real-world practice. Method: This study pioneers the systematic integration of translation norm theory into machine translation, proposing a norm-explicit LLM-based translation framework. It leverages prompt engineering to elicit multi-style translations from large language models and employs a multidimensional evaluation combining expert error analysis, user preference ranking, and automated metrics. Results: Evaluated on investor relations texts from 33 listed companies, the method consistently outperforms official human translations in human evaluation. Its core contributions are: (1) establishing a norm-driven translation paradigm; (2) empirically validating that norm-guided MT can surpass professional human translation; and (3) providing an interpretable, controllable pathway for business-oriented machine translation.
Resource-constrained freelance translators struggle to adopt state-of-the-art translation technologies due to computational, technical, and workflow-integration barriers. Method: This study proposes a lightweight, embeddable translation analytics framework tailored for individual workflows, systematically adapting industrial-grade automatic evaluation metrics (BLEU, chrF, TER, COMET) to freelance translation practice. Leveraging a real-world trilingual medical translation corpus, we introduce an interpretable evaluation paradigm designed for low-resource, multilingual, domain-specific settings. Results: Statistical validation via human–machine score correlation demonstrates that COMET—and to a lesser extent other metrics—exhibits significant agreement with expert human judgments in medical translation (p < 0.01). The framework delivers cost-effective, high-fidelity quality diagnostics, fine-grained error localization, and actionable feedback for iterative improvement. It bridges a critical gap between automated evaluation research and micro-level professional translation practice.
This study addresses low accuracy and inter-annotator inconsistency in human evaluation of machine translation (MT), stemming from annotator capability variation and task design bias. We propose a paradigm shift from traditional pointwise annotation to pairwise side-by-side (SxS) comparison, grounded in the Multidimensional Quality Metrics (MQM) framework. Through systematic comparison of MQM, SxS-MQM, and SxS relative ranking (RR), we empirically demonstrate—for the first time—that SxS significantly improves inter-annotator agreement (average 19.5–38.5% gain in error-label consistency) and cross-system error detection stability, especially for subtle semantic and stylistic deviations often missed by MQM. All settings preserve system-level ranking stability, while SxS-RR achieves the optimal trade-off between evaluation efficiency and reliability. We publicly release a triple-annotated dataset comprising 377 Chinese–English and 104 English–German sentence pairs, establishing a new benchmark and reproducible resource for MT evaluation.
In high-quality machine translation (MT) human evaluation, performance gains are often obscured by annotation noise. To address this, we propose a two-stage MQM-based collaborative re-annotation method: building upon initial annotations, it introduces iterative human review and collaborative refinement to jointly optimize primary annotations, peer annotations, and automatic predictions. This work is the first to integrate collaborative editing behavior modeling into the MQM framework. It significantly improves annotation consistency (+18.3%) and error detection rate (+24.7%), effectively recovering errors missed in the first round. Experiments demonstrate that re-annotation substantially enhances the reliability and stability of evaluation outcomes, enabling more accurate reflection of model quality improvements in assessment scores. The proposed approach establishes a scalable, reproducible paradigm for high-precision MT evaluation.
Existing post-editing methods for machine translation fail to fully harness the capabilities of large language models (LLMs). This paper proposes a novel LLM-based post-editing framework that integrates fine-grained MQM error annotations with LLMs: for the first time, MQM quality labels are incorporated as interpretable external feedback into both prompting and supervised fine-tuning of LLaMA-2, enabling error-driven, precise editing. The method comprises MQM annotation parsing, multilingual modeling (Chinese–English, English–German, English–Russian), and feedback-aware instruction tuning. Experiments demonstrate consistent improvements over baselines across TER, BLEU, and COMET metrics; human evaluation confirms substantial gains in translation quality, while fine-tuning significantly enhances the model’s efficiency in leveraging fine-grained feedback. The core contribution is a principled, interpretable, and learnable MQM–LLM collaborative post-editing paradigm.
This study investigates the quality differences between machine translation (MT) and human post-editing (PE) in domain-specific English-to-French translation tasks, as well as the influence of post-editors’ professional backgrounds on editing outcomes. The experiment compares three leading MT systems—DeepL, eTranslation, and Systran—and involves two groups of post-editors: linguists/translators and NLP experts. A fine-grained error taxonomy tailored to MT and PE is employed for manual evaluation. For the first time in a specialized translation setting, the interaction between MT system performance and editor background is jointly examined. Findings reveal that terminological accuracy and linguistic fluency are significantly affected by domain expertise, highlighting current limitations of MT in handling specialized language and underscoring the value of interdisciplinary post-editing teams.
Current automatic Machine Translation Quality Estimation (QE) systems lack reliability in real-world scenarios because their segment-level evaluations neglect critical dimensions such as discourse coherence, stylistic consistency, and rhetorical adequacy. Through theoretical analysis and empirical investigation, this study systematically uncovers structural limitations in QE—particularly concerning generalization capacity, data bias, overfitting, annotation noise, and the modeling of linguistic complexity—and, for the first time, identifies an inherent bottleneck rooted in the cognitive foundations of translation itself. These findings challenge the prevailing assumption that QE performance can be sufficiently improved merely by scaling up data or model capacity. The paper cautions against deploying existing QE systems as the sole basis for decision routing or bypassing human review in production environments and advocates redirecting future research toward automating human evaluation grounded in the Multidimensional Quality Metrics (MQM) framework.
This study investigates the development of students’ critical evaluation capabilities in AI-assisted translation pedagogy by engaging them in comparative assessments of outputs from general-purpose large language models and online machine translation systems. Participants performed post-editing tasks and justified their decisions using both automatic metrics (e.g., BLEU) and human evaluations focusing on fluency and accuracy. Findings reveal that students primarily based their judgments on multidimensional criteria—including terminological precision, linguistic naturalness, and anticipated editing effort—and often preferred translations that diverged from rankings suggested by automatic scores, demonstrating a nuanced, critical judgment that transcends quantitative metrics. This work challenges conventional paradigms of system evaluation by uncovering the complexity and pedagogical significance of learners’ subjective assessment logic in authentic instructional contexts.
This study addresses the insufficient reliability of existing automatic evaluation metrics in historical language scenarios such as classical Chinese-to-English translation, which hinders their utility in digital humanities research. The authors propose the first diagnostic framework based on minimal pairs to systematically assess both reference-dependent and reference-free metrics with respect to their sensitivity to translation errors and tolerance for legitimate variants. Experimental results reveal that all evaluated metrics exhibit blind spots, though MetricX24 demonstrates superior performance overall. By uncovering the limitations of current metrics in cross-cultural historical translation contexts, this work introduces a diagnostic approach tailored to scholarly applications and lays the groundwork for developing more robust and interpretable evaluation tools.
This study addresses underexamined translation errors in current machine translation benchmarks that may compromise the reliability and comparability of multilingual large language model (LLM) evaluations. It presents the first systematic quantification of the isolated impact of target-side translation errors on multilingual LLM assessment outcomes. The approach leverages an LLM-based evaluator to generate MQM-style error annotations, integrates the xCOMET-XXL quality estimation model, and employs controlled variable analysis while holding the correctness of English source prompts constant. Findings indicate that although automatically generated error annotations exhibit discrepancies compared to human judgments, translation errors nonetheless induce a statistically significant drop in model accuracy. This result validates the efficacy of automated error localization methods when applied to real-world translation benchmarks.