🤖 AI Summary
This study investigates how morphological alignment affects tokenization and downstream syntactic tasks—part-of-speech tagging, named entity recognition, and dependency parsing—in highly agglutinative languages like Telugu. We construct a Telugu segmentation dataset annotated with fine-grained morphological boundaries and comparatively evaluate Unigram and Byte-Pair Encoding (BPE) tokenizers in a multilingual setting, incorporating morphological alignment strategies and a novel morphology-aware hybrid tokenization method. Results show that Unigram substantially outperforms BPE; morphological alignment yields only moderate gains for syntactic tasks; the proposed hybrid tokenizer effectively mitigates BPE’s limitations; and intrinsic metrics—including Connectionist Temporal Classification (CTC) loss and Rényi entropy—exhibit no significant correlation with downstream performance. The core contribution lies in uncovering systematic interactions between tokenization paradigms and morphological modeling, and proposing a practical, morphology-informed hybrid tokenization strategy that balances linguistic fidelity with computational efficiency.
📝 Abstract
Prior work on language modeling showed conflicting findings about whether morphologically aligned approaches to tokenization improve performance, particularly for languages with complex morphology. To investigate this, we select a typologically diverse set of languages: Telugu (agglutinative), Hindi (primarily fusional with some agglutination), and English (fusional). We conduct a comprehensive evaluation of language models -- starting from tokenizer training and extending through the finetuning and downstream task evaluation. To account for the consistent performance differences observed across tokenizer variants, we focus on two key factors: morphological alignment and tokenization quality. To assess morphological alignment of tokenizers in Telugu, we create a dataset containing gold morpheme segmentations of 600 derivational and 7000 inflectional word forms.
Our experiments reveal that better morphological alignment correlates positively -- though moderately -- with performance in syntax-based tasks such as Parts-of-Speech tagging, Named Entity Recognition and Dependency Parsing. However, we also find that the tokenizer algorithm (Byte-pair Encoding vs. Unigram) plays a more significant role in influencing downstream performance than morphological alignment alone. Naive Unigram tokenizers outperform others across most settings, though hybrid tokenizers that incorporate morphological segmentation significantly improve performance within the BPE framework. In contrast, intrinsic metrics like Corpus Token Count (CTC) and Rényi entropy showed no correlation with downstream performance.