🤖 AI Summary
This study addresses the unclear impact of tokenizer selection on multilingual model performance across languages, particularly low-resource ones. To investigate this, we conduct large-scale controlled experiments by training 123 models using 54 distinct tokenizers. We introduce a systematic evaluation framework incorporating the bits-per-byte metric, Spearman correlation analysis, and attribute-based predictive models. Our findings reveal how tokenizer influence varies with data scale, demonstrating that low-resource languages exhibit substantially greater sensitivity to tokenizer choice. Furthermore, we propose a candidate filtering strategy based on intrinsic tokenizer properties, enabling precise prediction of downstream model performance rankings through specific metrics. This work provides actionable insights for optimizing tokenizer selection in multilingual modeling, bridging the gap between intrinsic tokenizer characteristics and extrinsic task performance.
📝 Abstract
Tokenizer choice affects multilingual language modeling, but vocabulary capacity is finite and vocabulary size is often constrained: improving representation for some languages often comes at the expense of others. We therefore ask whether tokenizer choice matters equally across languages, a question that the current literature leave unanswered. To this end, we train 123 language models spanning 54 tokenizers. In the main comparison, architecture, training corpus, training-token budget, and optimization are held fixed, so the models differ only in their tokenizer. We find that tokenizer choice matters more for languages with less language-model training data: across the 54 tokenizers, the standard deviation of a language's bits-per-byte (BPB) increases as its model training-data share decreases (Spearman rho = -0.52 over the 31 trained languages and -0.69 over the 28 written with word boundaries). Leaving a language out of tokenizer training raises its BPB in every language we study, and the penalty tends to be larger for languages with less language-model training data. Giving lower-resource languages a larger share of tokenizer-training data, however, does not unconditionally help those languages: both equal weighting and an allocation inverting the shares with respect to the language model training data increase their BPB, particularly when language-model training repeats data. Finally, which intrinsic tokenizer properties are associated with better BPB differs across languages, providing further evidence that what makes a good tokenizer depends on the language. We find that the metrics quantifying these properties can be successfully used to predict downstream models' pairwise BPB rankings, suggesting a practical strategy for screening tokenizer candidates before training language models.