Translation as Augmentation: Effect of Translated Data on Assessment of Difficulty

📅 2026-07-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of training effective text complexity assessment models for low-resource languages, which suffer from a scarcity of fine-grained difficulty-annotated data. To overcome this limitation, the authors propose a cross-lingual data augmentation strategy that leverages machine translation to transfer annotated corpora from high-resource languages to low-resource European languages. By integrating translated data with limited native annotations, they train a BERT-based regression model for text difficulty prediction. Experimental results demonstrate that incorporating translated data significantly enhances model performance under conditions of extreme native data scarcity, thereby providing the first systematic validation of the utility of translation-derived data as a complementary resource for text complexity assessment in low-resource linguistic settings.
📝 Abstract
Reliable Text Difficulty Assessment is a prerequisite for valid text simplification workflows and personalized learning applications. However, the development of robust assessment models is severely hindered by a critical bottleneck: the scarcity of expert-annotated corpora containing fine-grained difficulty levels (e.g., CEFR), particularly for lower-resource languages. This paper addresses this data scarcity problem in the context of a low-resource European language. We propose a cross-lingual data augmentation strategy that leverages machine translation to transfer labeled resources from high-resource languages to the target low-resource language. We train BERT-based regression models to predict difficulty scores and investigate whether synthetic, translated data can effectively supplement native training sets. Our experiments demonstrate that augmenting scarce native data with machine-translated corpora significantly improves the accuracy of difficulty estimation, offering a viable solution for languages lacking extensive expert annotations.
Problem

Research questions and friction points this paper is trying to address.

text difficulty assessment
low-resource languages
expert-annotated corpora
data scarcity
CEFR
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-lingual data augmentation
machine translation
text difficulty assessment
low-resource languages
BERT-based regression
🔎 Similar Papers
No similar papers found.