TermJudge: A Document-Level Metric Judging, Not Counting, Terminology in Machine Translation Evaluation

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing machine translation evaluation metrics in distinguishing terminology translation errors from legitimate variants, which introduces assessment bias. To overcome this, we propose TermJudge, a document-level terminology evaluation metric that pioneers a “judge rather than count” paradigm. Adopting an LLM-as-Judge framework, TermJudge implements a two-stage process: it first detects candidate terms via deterministic matching and subsequently leverages large language models to precisely differentiate valid variants from genuine errors within full-document contexts, thereby transcending the constraints of traditional rigid matching. Empirical results demonstrate that TermJudge ranks first in both system-level and segment-level meta-evaluations. Furthermore, this work validates that injecting terminology glossaries effectively eliminates translation errors and significantly enhances overall translation quality.
📝 Abstract
Existing automatic metrics for evaluating terminological use in machine translation (MT) penalise any divergence from a fixed reference, conflating translation errors with the valid terminological variation that human translators routinely produce. We introduce TermJudge, a document-level terminology metric that assigns an interpretable verdict to every term occurrence: glossary-conforming occurrences are settled deterministically, while divergences are assessed under a two-step LLM-as-judge procedure using the full document context: the first detects and labels terminology errors; the second sorts valid document-level variations from inconsistencies. Validated against expert error annotations and document-level human MQM scores, TermJudge ranks first in both system- and segment-level meta-evaluation, ahead of glossary-conformity and quality-estimation baselines. When applied to eight systems translating academic documents, under two prompting conditions, we observe that glossary injection improves terminology translation in all paired comparisons, by removing genuine errors rather than valid variation. TermJudge is released as open-source code.
Problem

Research questions and friction points this paper is trying to address.

Machine Translation Evaluation
Terminology Metrics
Document-Level Assessment
Terminological Variation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Terminology Evaluation
LLM-as-Judge
Document-Level Metric
Machine Translation
TermJudge
🔎 Similar Papers
No similar papers found.