๐ค AI Summary
This study investigates how subword segmentation strategies affect morphological similarity modeling in low-resource Kurmanji Kurdish. Method: We propose a bootstrapped BiLSTM-CRF morphological segmenter trained on minimal human annotations and integrate it with Word2Vec to generate word embeddings. We systematically compare word-level, morpheme-level, and Byte-Pair Encoding (BPE) segmentation under a multi-dimensional evaluation framework assessing similarity preservation, clustering quality, and semantic structure fidelity. Contribution/Results: We identify a critical coverage bias in existing evaluationsโBPE covers only 28.6% of test instances, leading to inflated performance estimates. In contrast, morpheme-level segmentation achieves 68.7% coverage and significantly outperforms BPE in semantic neighborhood quality and balanced representation of morphological complexity. These findings underscore the necessity of coverage-aware evaluation for low-resource NLP tasks, revealing that segmentation granularity and lexical coverage jointly determine embedding efficacy in morphologically rich, data-scarce languages.
๐ Abstract
We investigate tokenization strategies for Kurdish word embeddings by comparing word-level, morpheme-based, and BPE approaches on morphological similarity preservation tasks. We develop a BiLSTM-CRF morphological segmenter using bootstrapped training from minimal manual annotation and evaluate Word2Vec embeddings across comprehensive metrics including similarity preservation, clustering quality, and semantic organization. Our analysis reveals critical evaluation biases in tokenization comparison. While BPE initially appears superior in morphological similarity, it evaluates only 28.6% of test cases compared to 68.7% for morpheme model, creating artificial performance inflation. When assessed comprehensively, morpheme-based tokenization demonstrates superior embedding space organization, better semantic neighborhood structure, and more balanced coverage across morphological complexity levels. These findings highlight the importance of coverage-aware evaluation in low-resource language processing and offers different tokenization methods for low-resourced language processing.