🤖 AI Summary
Current automatic Machine Translation Quality Estimation (QE) systems lack reliability in real-world scenarios because their segment-level evaluations neglect critical dimensions such as discourse coherence, stylistic consistency, and rhetorical adequacy. Through theoretical analysis and empirical investigation, this study systematically uncovers structural limitations in QE—particularly concerning generalization capacity, data bias, overfitting, annotation noise, and the modeling of linguistic complexity—and, for the first time, identifies an inherent bottleneck rooted in the cognitive foundations of translation itself. These findings challenge the prevailing assumption that QE performance can be sufficiently improved merely by scaling up data or model capacity. The paper cautions against deploying existing QE systems as the sole basis for decision routing or bypassing human review in production environments and advocates redirecting future research toward automating human evaluation grounded in the Multidimensional Quality Metrics (MQM) framework.
📝 Abstract
Automation of Translation Quality Estimation (QE) has emerged as a widely discussed approach to managing translation quality at scale, and a growing number of tools and technologies have been released in pursuit of this goal. However, the proliferation of new QE systems has not always been accompanied by robust, transparent, and reproducible research and testing. This gap deserves critical scrutiny. This paper examines some fundamental limitations of the QE technology from both theoretical and empirical perspectives, arguing that current QE systems are structurally ill-equipped to serve as reliable standalone tools in real-world translation workflows. The reviewed evidence suggests that QE suffers from a range of interrelated and largely unresolved limitations. Most fundamentally, the evaluation of the quality of translation at the level of isolated segments is problematic because it tends to miss out on cohesion, coherence, and stylistic and rhetorical text features. In addition, empirical research documents several other limitations and flaws, including failure to generalize, systematic biases, overfitting and distribution collapse, performance gaps, error annotation challenges, and data scarcity. These are structural limitations arising from the complexity of human language and translation as a cognitive and communicative act - limitations that more data and better architectures have so far not overcome. Consequently, segment-level QE scores should not be used as a standalone basis for routing, release, or review bypass in production; we argue future work should focus on automating human evaluation grounded in MQM.