🤖 AI Summary
Universal Dependencies (UD) treebanks, designed primarily for native language data, inadequately capture syntactic errors characteristic of second-language (L2) Korean learners. Method: We construct and expand the first L2-oriented Korean UD treebank, adding 5,454 manually annotated sentences; systematically revise the Korean UD annotation guidelines for L2 phenomena; and propose error-aware data augmentation and fine-grained annotation strategies to bridge the gap between native UD standards and L2 linguistic reality. The treebank is formally aligned with the UD v2 framework. Contribution/Results: We conduct domain-adaptive fine-tuning and cross-domain evaluation on KoBERT, KLUE-RoBERTa, and XLM-R. Results show significant improvements in dependency parsing F1 scores on both in-domain and out-of-domain L2 test sets, demonstrating that purpose-built L2 treebanks critically enhance the robustness of morphosyntactic analysis for non-native Korean.
📝 Abstract
We expand the second language (L2) Korean Universal Dependencies (UD) treebank with 5,454 manually annotated sentences. The annotation guidelines are also revised to better align with the UD framework. Using this enhanced treebank, we fine-tune three Korean language models and evaluate their performance on in-domain and out-of-domain L2-Korean datasets. The results show that fine-tuning significantly improves their performance across various metrics, thus highlighting the importance of using well-tailored L2 datasets for fine-tuning first-language-based, general-purpose language models for the morphosyntactic analysis of L2 data.