Subword Tokenization Strategies for Kurdish Word Embeddings

๐Ÿ“… 2025-11-18
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study investigates how subword segmentation strategies affect morphological similarity modeling in low-resource Kurmanji Kurdish. Method: We propose a bootstrapped BiLSTM-CRF morphological segmenter trained on minimal human annotations and integrate it with Word2Vec to generate word embeddings. We systematically compare word-level, morpheme-level, and Byte-Pair Encoding (BPE) segmentation under a multi-dimensional evaluation framework assessing similarity preservation, clustering quality, and semantic structure fidelity. Contribution/Results: We identify a critical coverage bias in existing evaluationsโ€”BPE covers only 28.6% of test instances, leading to inflated performance estimates. In contrast, morpheme-level segmentation achieves 68.7% coverage and significantly outperforms BPE in semantic neighborhood quality and balanced representation of morphological complexity. These findings underscore the necessity of coverage-aware evaluation for low-resource NLP tasks, revealing that segmentation granularity and lexical coverage jointly determine embedding efficacy in morphologically rich, data-scarce languages.

Technology Category

Natural Language Processing: Lexical Semantics and MorphologyData Mining & Knowledge Management: Semantic WebMachine Learning: Evaluation and Analysis

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web data
๐Ÿ“ Abstract
We investigate tokenization strategies for Kurdish word embeddings by comparing word-level, morpheme-based, and BPE approaches on morphological similarity preservation tasks. We develop a BiLSTM-CRF morphological segmenter using bootstrapped training from minimal manual annotation and evaluate Word2Vec embeddings across comprehensive metrics including similarity preservation, clustering quality, and semantic organization. Our analysis reveals critical evaluation biases in tokenization comparison. While BPE initially appears superior in morphological similarity, it evaluates only 28.6% of test cases compared to 68.7% for morpheme model, creating artificial performance inflation. When assessed comprehensively, morpheme-based tokenization demonstrates superior embedding space organization, better semantic neighborhood structure, and more balanced coverage across morphological complexity levels. These findings highlight the importance of coverage-aware evaluation in low-resource language processing and offers different tokenization methods for low-resourced language processing.
Problem

Research questions and friction points this paper is trying to address.

Evaluating tokenization strategies for Kurdish word embeddings
Developing morphological segmenter with minimal manual annotation
Analyzing evaluation biases in tokenization method comparisons
Innovation

Methods, ideas, or system contributions that make the work stand out.

Developed BiLSTM-CRF morphological segmenter with bootstrapped training
Compared word-level, morpheme-based and BPE tokenization strategies
Revealed morpheme-based tokenization superior in embedding organization
๐Ÿ”Ž Similar Papers
No similar papers found.
University at Buffalo
A
Ali Salehi
Department of Linguistics, University at Buffalo, Buffalo NY, USA
C
Cassandra L. Jacobs
Department of Linguistics, University at Buffalo, Buffalo NY, USA