Rethinking Tokenization for Rich Morphology: The Dominance of Unigram over BPE and Morphological Alignment

📅 2025-08-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates how morphological alignment affects tokenization and downstream syntactic tasks—part-of-speech tagging, named entity recognition, and dependency parsing—in highly agglutinative languages like Telugu. We construct a Telugu segmentation dataset annotated with fine-grained morphological boundaries and comparatively evaluate Unigram and Byte-Pair Encoding (BPE) tokenizers in a multilingual setting, incorporating morphological alignment strategies and a novel morphology-aware hybrid tokenization method. Results show that Unigram substantially outperforms BPE; morphological alignment yields only moderate gains for syntactic tasks; the proposed hybrid tokenizer effectively mitigates BPE’s limitations; and intrinsic metrics—including Connectionist Temporal Classification (CTC) loss and Rényi entropy—exhibit no significant correlation with downstream performance. The core contribution lies in uncovering systematic interactions between tokenization paradigms and morphological modeling, and proposing a practical, morphology-informed hybrid tokenization strategy that balances linguistic fidelity with computational efficiency.

Technology Category

Natural Language Processing: Lexical Semantics and MorphologyMachine Learning: Mixture of Experts (MoE)Computer Vision: Language and Vision

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Prior work on language modeling showed conflicting findings about whether morphologically aligned approaches to tokenization improve performance, particularly for languages with complex morphology. To investigate this, we select a typologically diverse set of languages: Telugu (agglutinative), Hindi (primarily fusional with some agglutination), and English (fusional). We conduct a comprehensive evaluation of language models -- starting from tokenizer training and extending through the finetuning and downstream task evaluation. To account for the consistent performance differences observed across tokenizer variants, we focus on two key factors: morphological alignment and tokenization quality. To assess morphological alignment of tokenizers in Telugu, we create a dataset containing gold morpheme segmentations of 600 derivational and 7000 inflectional word forms. Our experiments reveal that better morphological alignment correlates positively -- though moderately -- with performance in syntax-based tasks such as Parts-of-Speech tagging, Named Entity Recognition and Dependency Parsing. However, we also find that the tokenizer algorithm (Byte-pair Encoding vs. Unigram) plays a more significant role in influencing downstream performance than morphological alignment alone. Naive Unigram tokenizers outperform others across most settings, though hybrid tokenizers that incorporate morphological segmentation significantly improve performance within the BPE framework. In contrast, intrinsic metrics like Corpus Token Count (CTC) and Rényi entropy showed no correlation with downstream performance.
Problem

Research questions and friction points this paper is trying to address.

Evaluating morphological alignment impact on tokenization performance
Comparing Unigram and BPE tokenizers for diverse languages
Assessing tokenization quality in syntax-based NLP tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unigram tokenizers outperform BPE tokenizers
Hybrid tokenizers enhance BPE with morphology
Morphological alignment moderately aids syntax tasks
🔎 Similar Papers
2024-06-21arXiv.orgCitations: 0
💼 Related Jobs
No related jobs found.
Saketh Reddy Vemula
Saketh Reddy Vemula
IIIT Hyderabad
Natural Language ProcessingComputational Linguistics
D
Dipti Mishra Sharma
IIIT Hyderabad, India
P
Parameswari Krishnamurthy
IIIT Hyderabad, India