ThaiTrees: Thai Syntactic Dependency Trees Across Domains

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为了解决泰语文本缺乏大规模自动解析语料库的问题,本文通过开发可重复使用的处理管道,创建了一个3.42亿词的多领域泰语句法依存树库。
📝 Abstract
Studying syntactic patterns in naturally occurring language requires a large parsed corpus, but manual annotation is costly and difficult to scale. Thai has a manually annotated dependency treebank for training and evaluating parsers, but lacks a large automatically parsed corpus for quantitative syntactic research. We present ThaiTrees, a 342M-token corpus drawn from news, Wikipedia, spoken transcripts, and social media. We develop a reproducible pipeline for cleaning, processing, and parsing Thai text under the Universal Dependencies framework. The resulting corpus makes grammatical relations searchable and supports the study of syntactic distributions. We release a frequency lexicon and CoNLL-U parses in machine-readable formats suitable for both AI-assisted and conventional programmatic analysis.
Problem

Research questions and friction points this paper is trying to address.

Thai
syntactic dependency
parsed corpus
quantitative syntactic research
Innovation

Methods, ideas, or system contributions that make the work stand out.

ThaiTrees
Universal Dependencies
automatically parsed corpus
syntactic distribution
machine-readable format
🔎 Similar Papers
No similar papers found.
A
Attapol T. Rutherford
Department of Linguistics, Chulalongkorn University
P
Papatchol Thientong
Department of Linguistics, Chulalongkorn University