construct ud treebank

Designs and builds Universal Dependencies–style morphologically annotated corpora (UD treebanks), producing sentence-level tokenization, lemmas, POS and morphological feature annotations, and labeled dependency relations serialized in CoNLL‑U or equivalent formats. This competence includes producing annotation guidelines and conversion tools, validating and versioning the dataset, and creating deterministic train/dev/test splits for reproducible evaluation.

constructudtreebank

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.71
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenge of converting the large-scale, multi-genre Czech Prague Dependency Treebank (PDT-C) into Universal Dependencies (UD) format while preserving high annotation quality. The work tackles systematic discrepancies between the two frameworks in syntactic structure, part-of-speech tagging, and granularity of dependency relations. By designing a fine-grained cross-scheme mapping strategy combined with multi-layer linguistic alignment and dependency topology adjustments, the authors achieve the first large-scale conversion from PDT-C to UD. The resulting treebank, UD_Czech-PDTC, is currently the largest and most genre-diverse Czech UD resource, more than doubling the size of the original PDT. This significantly advances standardization and cross-domain coverage for low-resource language treebanks, providing a high-quality foundational resource for multilingual natural language processing.

annotation conversiondependency parsingPrague Dependency Treebank

This study addresses the limitations of the Universal Dependencies (UD) framework in capturing syntactic phenomena that span multiple speaker turns in spoken language, such as collaborative utterance construction, question–answer interactions, and backchannel responses. To overcome these challenges, the paper proposes two annotation schemes: one based on turn segmentation and another employing a unified dependency structure that permits cross-turn dependencies. It further introduces novel strategies for handling constituent promotion in reformulations, repairs, and incomplete phrases. The work offers the first systematic delineation and differentiation of co-construction, reformulation, and repair in spoken discourse, transcending the traditional single-turn syntactic analysis paradigm. As a result, it establishes the first UD-compliant annotation guidelines explicitly supporting cross-turn dependencies, whose feasibility and effectiveness are demonstrated through application to a real-world spoken treebank.

annotation guidelinescoconstructionsspoken language

UD-KSL Treebank v1.3: A semi-automated framework for aligning XPOS-extracted units with UPOS tags

Jun 10, 2025
HS
Hakyung Sung
🏛️ University of Oregon | University of Illinois Chicago | Konkuk University | Chung-Ang University | Yale University

This study addresses the challenge of morphosyntactic structure identification and alignment from XPOS (language-specific part-of-speech) sequences to UPOS (universal part-of-speech) tags in second-language Korean (L2-Korean) Universal Dependencies (UD) annotation. Method: We propose the first fine-grained, structured XPOS→UPOS cross-layer alignment framework, integrating rule-based heuristics with statistical models via fine-tuning spaCy and UDPipe for semi-automatic alignment. We further augment the L2-Korean corpus with 2,998 newly annotated argumentative essays. Contribution/Results: Our work establishes the first explicit, structured mapping between XPOS and UPOS, substantially improving multi-layer annotation consistency. In low-resource settings, it significantly enhances both morphosyntactic analysis and dependency parsing accuracy—demonstrating the efficacy and generalizability of cross-layer alignment for downstream NLP tasks.

Aligns XPOS sequences with UPOS tags for L2 KoreanEvaluates impact of XPOS-UPOS alignments on NLP modelsExpands L2-Korean corpus with annotated argumentative essays

Second language Korean Universal Dependency treebank v1.2: Focus on data augmentation and annotation scheme refinement

Mar 18, 2025
HS
Hakyung Sung
🏛️ University of Oregon | University of Illinois Chicago

Universal Dependencies (UD) treebanks, designed primarily for native language data, inadequately capture syntactic errors characteristic of second-language (L2) Korean learners. Method: We construct and expand the first L2-oriented Korean UD treebank, adding 5,454 manually annotated sentences; systematically revise the Korean UD annotation guidelines for L2 phenomena; and propose error-aware data augmentation and fine-grained annotation strategies to bridge the gap between native UD standards and L2 linguistic reality. The treebank is formally aligned with the UD v2 framework. Contribution/Results: We conduct domain-adaptive fine-tuning and cross-domain evaluation on KoBERT, KLUE-RoBERTa, and XLM-R. Results show significant improvements in dependency parsing F1 scores on both in-domain and out-of-domain L2 test sets, demonstrating that purpose-built L2 treebanks critically enhance the robustness of morphosyntactic analysis for non-native Korean.

Evaluate fine-tuned models on L2-Korean datasetsExpand L2 Korean UD treebank with annotated sentencesRefine annotation guidelines to align with UD framework

A Systematic Comparison of Syntactic Representations of Dependency Parsing

May 29, 2017
GW
Guillaume Wisniewski
🏛️ Univ. Paris-Sud | Université Paris-Saclay | University of Copenhagen

This study systematically investigates how dependency annotation schemes affect the performance of transition-based parsers. Method: Addressing language-specific non-canonical structures in Universal Dependencies (UD) treebanks, we design standardization transformation rules and comparatively evaluate parser performance—measured by LAS and UAS—under both original and standardized annotations within a unified, multilingual evaluation framework. Contribution/Results: We empirically demonstrate, for the first time, that annotation standardization does not universally improve parsing accuracy. Crucially, we reveal that linguistic typological features significantly moderate the effectiveness of annotation schemes: for certain languages, the original non-standard annotations yield higher accuracy than standardized ones. This finding challenges the implicit assumption that standardization is inherently optimal and underscores the necessity of considering language-specific syntactic properties when selecting or designing syntactic representations.

Compare parser performance across annotation schemes.Convert syntactic constructions to standard representations.Evaluate parsing performance across multiple languages.

Latest Papers

What's happening recently
View more

This work addresses the lack of a unified, coherent, and high-quality Czech treebank that systematically covers cross-sentential semantic phenomena such as coreference and discourse relations. Developed over nearly three decades within the framework of Prague Dependency Grammar, the resulting corpus spans multiple genres and comprises approximately four million words with multi-layer annotations. It is the first to integrate deep semantic representations, coreference resolution, and discourse relation annotation in a fully consistent and uniform manner across the entire dataset. The treebank has become a widely used international resource, supporting both traditional and modern NLP tool evaluation as well as cross-formalism conversion studies. All annotated data and associated parsers have been publicly released under open-source licenses.

annotation schemecoreferencediscourse relations

Constituency Structure over Eojeol in Korean Treebanks

Dec 27, 2025
JP
Jungyeul Park
🏛️ KAIST | Anyang University

Korean phrase-structure treebanks have long faced a terminal-unit selection dilemma: using morphemes as terminals conflates intra-word morphology with syntactic structure and impedes alignment with eojeol-based dependency treebanks. This paper proposes a phrase-structure representation with eojeol (orthographic word) as the terminal unit. We formally establish, for the first time, phrase-structure equivalence between the Sejong and Penn Korean treebanks at the eojeol level. We design a two-tier annotation framework: an upper tier encoding pure syntactic constituency, and a lower tier independently representing morphological and part-of-speech information. Additionally, we introduce an explicit constituency–dependency mapping model. Our approach achieves strict decoupling of syntax and morphology, enabling unified cross-treebank representation, high-fidelity parsing, paradigm conversion (e.g., constituency ↔ dependency), and joint modeling across heterogeneous linguistic resources.

Addresses terminal unit choice in Korean constituency treebanksEnables cross-treebank comparison and constituency-dependency conversionProposes eojeol-based representation to separate morphology from syntax

This work addresses the challenges posed by morphologically rich languages such as Czech, where extensive inflectional and derivational morphology leads to unwieldy lexicon sizes and difficulties in ensuring annotation consistency. To tackle these issues, the authors propose MorfFlex, an architecture that explicitly models inflectional and derivational rules through a system of structured pattern rules operating on <word form, lemma, part-of-speech> triples. By leveraging hand-curated source files and transformation scripts, MorfFlex compresses massive word-form inventories into a compact, manageable lexical resource. The resulting MorfFlex CZ dictionary encompasses over 100 million word forms and one million lemmas, drastically reducing storage requirements while supporting consistent annotation in the Prague Dependency Treebank and enabling high-quality NLP tools such as MorphoDiTa.

derivationinflectionlexical resource

为了解决泰语文本缺乏大规模自动解析语料库的问题,本文通过开发可重复使用的处理管道,创建了一个3.42亿词的多领域泰语句法依存树库。

parsed corpusquantitative syntactic researchsyntactic dependency

Hot Scholars

TE

Toqeer Ehsan

Teknologian tutkimuskeskus VTT Oy
Natural Language ProcessingDeep LearningArtificial Intelligence
JZ

Jiapeng Zhu

East China Normal University
Reinforcement LearningLarge Language ModelGraph Representation Learning
XY

Xiulin Yang

King Abdullah University of Science and Technology
Water splittingCO2 electroreductionClean energy
HA

Hassan Alhuzali

Assistant Professor @ Umm Al-Qura University
natural language processingcultural awareness of LLMsMental Healthaffective computing