๐ค AI Summary
This work addresses the challenge of achieving language-specific tokenization without modifying the vocabulary of pretrained models, thereby enhancing downstream task performance. The authors propose LangMAP, an extension of UnigramLM to multilingual settings, which jointly trains a shared vocabulary with language identifiers. Notably, LangMAP enables adaptive, language-specific tokenization during inference without requiring explicit language tagsโa first in the field. The method is compatible with both training-from-scratch and fine-tuning paradigms. Evaluated across nine natural languages and nine programming languages, LangMAP significantly improves morphological boundary detection and alignment with abstract syntax tree (AST) leaf nodes. Consistent gains on the MultiBLiMP grammaticality judgment benchmark further demonstrate its effectiveness and broad applicability.
๐ Abstract
Language-specific tokenizers improve tokenization quality and the downstream performance of models on those languages. However, using such a tokenizer comes at a cost: either a new model must be trained from scratch, or the vocabulary of an existing pretrained model must be adapted. We propose Language-adaptive Maximum a Posteriori (LangMAP) Tokenization, a tokenization scheme that extends the UnigramLM algorithm to the multilingual setting, producing language-specific tokenization from a single shared vocabulary. Notably, LangMAP can be used when training a multilingual language model from scratch or to adapt a pretrained model's tokenizer to individual languages without changing its vocabulary. While language labels are required at training time, a key feature of the algorithm is that it then performs language-specific tokenization at inference without knowledge of the input's language. Across 14 open-source tokenizers, 9 natural languages, and 9 programming languages, LangMAP improves morphological boundary alignment and, for all coding languages tested, alignment with abstract syntax tree (AST) leaf boundaries. In fine-tuning experiments, results are mixed: LangMAP improves target-language grammatical acceptability (MultiBLiMP) on the languages tested; its benefits are less consistent on knowledge-related tasks (Global-PIQA, Belebele).