Score
Designs, builds, and evaluates methods for mapping or aligning discrete and soft token representations and tokenizations—including learning soft-token embeddings, probability-weighted vocabulary mixtures, and tokenization-to-tokenization mappings—so that tokens from different vocabularies or tokenizers correspond to each other. Analyzes and optimizes these mappings by minimizing representation distance and preserving shared semantic structure across languages or alternative tokenizations.
This paper addresses the lack of theoretical foundations for tokenization in natural language processing (NLP), systematically investigating its impact on the statistical estimation consistency of language models. While prior work relies predominantly on empirical analysis, we introduce the first unified formal framework grounded in the category of random mappings to rigorously characterize the modeling essence of tokenizers. Our key contributions are: (1) necessary and sufficient conditions for tokenizers to preserve statistical estimation consistency; (2) a four-dimensional theoretical analysis framework—covering inconsistency, ambiguity, finiteness, and sequentiality; and (3) principled, verifiable tokenizer design criteria derived from the integration of category theory, statistical learning theory, and formal language theory. This work establishes the first rigorous mathematical foundation for representation reliability in neural language modeling.
Current multilingual tokenizers suffer from misaligned cross-lingual vocabularies, causing semantically equivalent words—e.g., English “I eat rice” and Hausa “Ina cin shinkafa”—to be mapped to distinct embeddings, severely hindering cross-lingual transfer for low-resource languages. To address this, we propose a parallel tokenizer framework: first training monolingual tokenizers independently, then aligning their vocabularies at the lexical level using bilingual dictionaries to enable semantically consistent cross-lingual subword sharing; finally constructing a unified, frequency-balanced shared semantic space. This work is the first to systematically reformulate vocabulary design in cross-lingual pretraining, enabling end-to-end Transformer pretraining. Pretrained on 13 low-resource languages, our model significantly outperforms mBERT and XLM-R on downstream tasks—including sentiment analysis and hate speech detection—demonstrating that vocabulary alignment is fundamental to cross-lingual generalization.
This work addresses the inconsistency in reasoning exhibited by multilingual large language models when presented with semantically equivalent prompts in different languages, a problem often exacerbated by discrete tokenization. To mitigate this issue, the authors propose SOLAR, a novel approach that introduces soft tokens—probabilistic weighted combinations of lexical embeddings—as continuous semantic representations during supervised fine-tuning. By aligning multilingual soft token spaces through English as a pivot, SOLAR reduces reliance on specific vocabularies or writing systems while preserving cross-lingually shared reasoning structures. Experimental results demonstrate that SOLAR significantly outperforms baseline methods across four multilingual reasoning benchmarks, achieving gains of up to 17.7 percentage points, with particularly pronounced improvements for low-resource languages. The method also substantially enhances cross-lingual representational similarity.
Traditional Byte-Pair Encoding (BPE) tokenization introduces token redundancy in low-resource languages, degrading the performance of small-scale models. Method: This paper proposes a BPE configuration method integrating hyperparameter optimization and compressed sensing. It systematically searches key BPE hyperparameters—including vocabulary size and merge iterations—and jointly evaluates configurations using intrinsic metrics (e.g., token count) and extrinsic task performance (generation and classification). Contribution/Results: The study provides the first empirical evidence that BPE configuration significantly impacts multilingual modeling for low-resource languages. Experiments across diverse languages and model scales show that optimal configurations reduce token counts by 12.7% on average and improve downstream task accuracy by 1.8–3.4 percentage points for small models. These gains substantially enhance modeling efficiency and generalization capability in low-resource settings.
This work addresses the low efficiency of key information identification and semantic retrieval in text. We first discover that large language model (LLM) text embeddings naturally align with salient input tokens in the latent space—a phenomenon empirically validated across diverse model architectures, training paradigms, and embedding methods. Leveraging this insight, we propose a principal-component-guided alignment enhancement method that decomposes the embedding geometry and explicitly steers representations toward critical tokens. Building upon this, we design a lightweight sparse retrieval paradigm that retains only ~20% of token embeddings while achieving 80% of dense retrieval performance. Experiments across eight mainstream LLM embedders confirm the robustness of the alignment mechanism. Our findings provide an interpretable foundation for sparse retrieval and instruction-tuned embeddings, and advance the understanding of the intrinsic nature of semantic relevance.
本文研究了语言建模中联合优化分词的效果,通过比较无分词器方法(如SSLMs和H-Nets)与固定分词器,发现联合优化改变了词汇结构,提高了模型性能。
This study investigates how subword segmentation strategies affect morphological similarity modeling in low-resource Kurmanji Kurdish. Method: We propose a bootstrapped BiLSTM-CRF morphological segmenter trained on minimal human annotations and integrate it with Word2Vec to generate word embeddings. We systematically compare word-level, morpheme-level, and Byte-Pair Encoding (BPE) segmentation under a multi-dimensional evaluation framework assessing similarity preservation, clustering quality, and semantic structure fidelity. Contribution/Results: We identify a critical coverage bias in existing evaluations—BPE covers only 28.6% of test instances, leading to inflated performance estimates. In contrast, morpheme-level segmentation achieves 68.7% coverage and significantly outperforms BPE in semantic neighborhood quality and balanced representation of morphological complexity. These findings underscore the necessity of coverage-aware evaluation for low-resource NLP tasks, revealing that segmentation granularity and lexical coverage jointly determine embedding efficacy in morphologically rich, data-scarce languages.
This study systematically investigates the intrinsic semantic and syntactic properties of mainstream word embedding methods—such as Word2Vec and GloVe—and their performance disparities across diverse natural language processing tasks. By establishing a unified evaluation framework that integrates publicly available pretrained embeddings with standard benchmark datasets, the work conducts empirical comparisons on canonical tasks including semantic similarity and analogical reasoning. The findings delineate the performance boundaries and optimal application scenarios for each embedding model, offering practitioners reliable guidance for model selection in real-world settings. Furthermore, the analysis deepens the understanding of the inherent limitations of static word representations, highlighting critical constraints in capturing contextual and compositional linguistic phenomena.
This study systematically evaluates bilingual lexicon induction (BLI) as a metric for assessing embedding space alignment quality, identifying its validity and limitations. Addressing diverse language pairs—including high- and low-resource settings and typologically distinct language families—the authors compare BLI performance across traditional linear alignment methods, multilingual pretrained models (e.g., mBERT, XLM-R), and novel compositional alignment strategies. They propose a stem-based BLI approach to improve matching accuracy for inflectional languages and introduce a more robust lexical pruning mechanism. Results show that compositional methods generally achieve the best BLI performance, while multilingual models exhibit pronounced advantages for low-resource languages. Crucially, standard BLI is found to be sensitive to vocabulary coverage bias and morphological noise, thus failing to reliably reflect true alignment quality. The work establishes a more rigorous methodological framework for evaluating embedding alignment, supported by extensive empirical analysis across linguistic settings.
This study addresses the unclear impact of tokenizer selection on multilingual model performance across languages, particularly low-resource ones. To investigate this, we conduct large-scale controlled experiments by training 123 models using 54 distinct tokenizers. We introduce a systematic evaluation framework incorporating the bits-per-byte metric, Spearman correlation analysis, and attribute-based predictive models. Our findings reveal how tokenizer influence varies with data scale, demonstrating that low-resource languages exhibit substantially greater sensitivity to tokenizer choice. Furthermore, we propose a candidate filtering strategy based on intrinsic tokenizer properties, enabling precise prediction of downstream model performance rankings through specific metrics. This work provides actionable insights for optimizing tokenizer selection in multilingual modeling, bridging the gap between intrinsic tokenizer characteristics and extrinsic task performance.