🤖 AI Summary
This study addresses the limitations of conventional tokenizers, where optimizing solely for compression rate leads to uneven multilingual capacity allocation and constrained representation quality. We propose the LCT framework, which decouples structure discovery from vocabulary construction. LCT introduces a morphology-driven latent core discovery mechanism that integrates minimum description length, entropy boundary signals, and morphosyntactic constraints to mine linguistic structures. By identifying reusable linguistic units, it constructs a cross-lingual shared vocabulary, demonstrating that compression rate is not the sole determinant of representation quality. Experiments across 104 languages show that LCT effectively reduces tokenization fertility and significantly improves MorphScore, yielding average performance gains of up to 2.00 points on downstream task benchmarks.
📝 Abstract
Tokenizers are commonly optimized for compression, but a compact vocabulary does not necessarily distribute its capacity evenly across languages. We introduce the Latent Core Tokenizer (LCT), a language-agnostic approach that separates structural discovery from vocabulary construction. LCT uses Minimum Description Length, entropy-based boundary signals, and morphotactic constraints to identify reusable linguistic units before constructing a shared vocabulary. Across 104 languages with a 200K-token vocabulary, LCT achieves lower fertility and higher MorphScore than BPE, Unigram, and parity-aware BPE, while maintaining comparable cross-lingual disparity in tokenization cost. Across four multilingual downstream benchmarks, LCT improves aggregate score by 1.48, 1.83, and 2.00 points over BPE, Unigram, and parity-aware BPE, respectively. Our findings show that compression alone does not predict representation quality and highlight the importance of morphology-driven structural discovery and how frequency is used to allocate the final vocabulary across languages.