Score
Designs and builds Universal Dependencies–style morphologically annotated corpora (UD treebanks), producing sentence-level tokenization, lemmas, POS and morphological feature annotations, and labeled dependency relations serialized in CoNLL‑U or equivalent formats. This competence includes producing annotation guidelines and conversion tools, validating and versioning the dataset, and creating deterministic train/dev/test splits for reproducible evaluation.
This study addresses the challenge of converting the large-scale, multi-genre Czech Prague Dependency Treebank (PDT-C) into Universal Dependencies (UD) format while preserving high annotation quality. The work tackles systematic discrepancies between the two frameworks in syntactic structure, part-of-speech tagging, and granularity of dependency relations. By designing a fine-grained cross-scheme mapping strategy combined with multi-layer linguistic alignment and dependency topology adjustments, the authors achieve the first large-scale conversion from PDT-C to UD. The resulting treebank, UD_Czech-PDTC, is currently the largest and most genre-diverse Czech UD resource, more than doubling the size of the original PDT. This significantly advances standardization and cross-domain coverage for low-resource language treebanks, providing a high-quality foundational resource for multilingual natural language processing.
This study addresses the limitations of the Universal Dependencies (UD) framework in capturing syntactic phenomena that span multiple speaker turns in spoken language, such as collaborative utterance construction, question–answer interactions, and backchannel responses. To overcome these challenges, the paper proposes two annotation schemes: one based on turn segmentation and another employing a unified dependency structure that permits cross-turn dependencies. It further introduces novel strategies for handling constituent promotion in reformulations, repairs, and incomplete phrases. The work offers the first systematic delineation and differentiation of co-construction, reformulation, and repair in spoken discourse, transcending the traditional single-turn syntactic analysis paradigm. As a result, it establishes the first UD-compliant annotation guidelines explicitly supporting cross-turn dependencies, whose feasibility and effectiveness are demonstrated through application to a real-world spoken treebank.
This study addresses the challenge of morphosyntactic structure identification and alignment from XPOS (language-specific part-of-speech) sequences to UPOS (universal part-of-speech) tags in second-language Korean (L2-Korean) Universal Dependencies (UD) annotation. Method: We propose the first fine-grained, structured XPOS→UPOS cross-layer alignment framework, integrating rule-based heuristics with statistical models via fine-tuning spaCy and UDPipe for semi-automatic alignment. We further augment the L2-Korean corpus with 2,998 newly annotated argumentative essays. Contribution/Results: Our work establishes the first explicit, structured mapping between XPOS and UPOS, substantially improving multi-layer annotation consistency. In low-resource settings, it significantly enhances both morphosyntactic analysis and dependency parsing accuracy—demonstrating the efficacy and generalizability of cross-layer alignment for downstream NLP tasks.
Universal Dependencies (UD) treebanks, designed primarily for native language data, inadequately capture syntactic errors characteristic of second-language (L2) Korean learners. Method: We construct and expand the first L2-oriented Korean UD treebank, adding 5,454 manually annotated sentences; systematically revise the Korean UD annotation guidelines for L2 phenomena; and propose error-aware data augmentation and fine-grained annotation strategies to bridge the gap between native UD standards and L2 linguistic reality. The treebank is formally aligned with the UD v2 framework. Contribution/Results: We conduct domain-adaptive fine-tuning and cross-domain evaluation on KoBERT, KLUE-RoBERTa, and XLM-R. Results show significant improvements in dependency parsing F1 scores on both in-domain and out-of-domain L2 test sets, demonstrating that purpose-built L2 treebanks critically enhance the robustness of morphosyntactic analysis for non-native Korean.
This study systematically investigates how dependency annotation schemes affect the performance of transition-based parsers. Method: Addressing language-specific non-canonical structures in Universal Dependencies (UD) treebanks, we design standardization transformation rules and comparatively evaluate parser performance—measured by LAS and UAS—under both original and standardized annotations within a unified, multilingual evaluation framework. Contribution/Results: We empirically demonstrate, for the first time, that annotation standardization does not universally improve parsing accuracy. Crucially, we reveal that linguistic typological features significantly moderate the effectiveness of annotation schemes: for certain languages, the original non-standard annotations yield higher accuracy than standardized ones. This finding challenges the implicit assumption that standardization is inherently optimal and underscores the necessity of considering language-specific syntactic properties when selecting or designing syntactic representations.
This work addresses the lack of a unified, coherent, and high-quality Czech treebank that systematically covers cross-sentential semantic phenomena such as coreference and discourse relations. Developed over nearly three decades within the framework of Prague Dependency Grammar, the resulting corpus spans multiple genres and comprises approximately four million words with multi-layer annotations. It is the first to integrate deep semantic representations, coreference resolution, and discourse relation annotation in a fully consistent and uniform manner across the entire dataset. The treebank has become a widely used international resource, supporting both traditional and modern NLP tool evaluation as well as cross-formalism conversion studies. All annotated data and associated parsers have been publicly released under open-source licenses.
Korean phrase-structure treebanks have long faced a terminal-unit selection dilemma: using morphemes as terminals conflates intra-word morphology with syntactic structure and impedes alignment with eojeol-based dependency treebanks. This paper proposes a phrase-structure representation with eojeol (orthographic word) as the terminal unit. We formally establish, for the first time, phrase-structure equivalence between the Sejong and Penn Korean treebanks at the eojeol level. We design a two-tier annotation framework: an upper tier encoding pure syntactic constituency, and a lower tier independently representing morphological and part-of-speech information. Additionally, we introduce an explicit constituency–dependency mapping model. Our approach achieves strict decoupling of syntax and morphology, enabling unified cross-treebank representation, high-fidelity parsing, paradigm conversion (e.g., constituency ↔ dependency), and joint modeling across heterogeneous linguistic resources.
This work addresses the challenges posed by morphologically rich languages such as Czech, where extensive inflectional and derivational morphology leads to unwieldy lexicon sizes and difficulties in ensuring annotation consistency. To tackle these issues, the authors propose MorfFlex, an architecture that explicitly models inflectional and derivational rules through a system of structured pattern rules operating on <word form, lemma, part-of-speech> triples. By leveraging hand-curated source files and transformation scripts, MorfFlex compresses massive word-form inventories into a compact, manageable lexical resource. The resulting MorfFlex CZ dictionary encompasses over 100 million word forms and one million lemmas, drastically reducing storage requirements while supporting consistent annotation in the Prague Dependency Treebank and enabling high-quality NLP tools such as MorphoDiTa.
针对中世纪文献拉丁语处理工具不足的问题,通过迭代纠正标注方法生成领域内训练数据,提高了解析器性能。
为了解决泰语文本缺乏大规模自动解析语料库的问题,本文通过开发可重复使用的处理管道,创建了一个3.42亿词的多领域泰语句法依存树库。