๐ค AI Summary
This study addresses the persistent challenge of machine translation in accurately conveying semantics in culturally rich texts, a task where current large language models (LLMs) exhibit notable deficiencies. Using *Dream of the Red Chamber* as a cultural lens, the authors construct the first bilingual ChineseโJapanese dataset specifically designed for culturally loaded translation. They systematically evaluate mainstream LLMs and identify three core challenges: inadequate task modeling, poor alignment between human judgments and model outputs, and limited reliability of existing automatic evaluation metrics. To this end, the work proposes an integrated evaluation framework grounded in cultural context, combining a rigorous human assessment protocol, analysis of automatic metrics, and fine-grained cultural annotation. Their findings reveal a significant performance gap among state-of-the-art models on culturally laden content and demonstrate that conventional evaluation methods fail to reliably assess translation quality in such contexts.
๐ Abstract
Culturally loaded translation poses unique challenges for machine translation (MT), as meanings are deeply embedded in socio-cultural contexts beyond surface linguistic forms. Although large language models (LLMs) have enabled MT systems to achieve human-like quality in many scenarios, their ability to handle culturally loaded expressions remains underexplored. In this study, we systematically investigate the challenges posed by culturally loaded translation in LLM-based MT systems. We construct a Chinese-Japanese bilingual dataset from the culturally representative corpus Dream of the Red Chamber, containing 500 segments across diverse cultural categories. Using a comprehensive evaluation protocol, we reveal three main challenges: (1) task challenges, where frontier LLMs exhibit notable performance gaps and struggle with culturally loaded content; (2) human evaluation challenges, where evaluator backgrounds lead to substantial disagreement in translation judgments; and (3) automatic evaluation challenges, where widely used metrics fail to reliably assess translation quality for this task. These findings may offer valuable insights for culture-oriented translation research in both computational science and linguistics.