token alignment

Designs, builds, and evaluates methods for mapping or aligning discrete and soft token representations and tokenizations—including learning soft-token embeddings, probability-weighted vocabulary mixtures, and tokenization-to-tokenization mappings—so that tokens from different vocabularies or tokenizers correspond to each other. Analyzes and optimizes these mappings by minimizing representation distance and preserving shared semantic structure across languages or alternative tokenizations.

tokenalignment

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.54
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

The Foundations of Tokenization: Statistical and Computational Concerns

Jul 16, 2024
JL
Juan Luis Gastaldi
🏛️ ETH Zürich | City University of New York

This paper addresses the lack of theoretical foundations for tokenization in natural language processing (NLP), systematically investigating its impact on the statistical estimation consistency of language models. While prior work relies predominantly on empirical analysis, we introduce the first unified formal framework grounded in the category of random mappings to rigorously characterize the modeling essence of tokenizers. Our key contributions are: (1) necessary and sufficient conditions for tokenizers to preserve statistical estimation consistency; (2) a four-dimensional theoretical analysis framework—covering inconsistency, ambiguity, finiteness, and sequentiality; and (3) principled, verifiable tokenizer design criteria derived from the integration of category theory, statistical learning theory, and formal language theory. This work establishes the first rigorous mathematical foundation for representation reliability in neural language modeling.

Addressing statistical and computational concerns in tokenizer designAnalyzing tokenization's impact on language model consistencyUnderstanding theoretical foundations of tokenization in NLP

Parallel Tokenizers: Rethinking Vocabulary Design for Cross-Lingual Transfer

Oct 07, 2025
MD
Muhammad Dehan Al Kautsar
🏛️ Mohamed bin Zayed University of Artificial Intelligence

Current multilingual tokenizers suffer from misaligned cross-lingual vocabularies, causing semantically equivalent words—e.g., English “I eat rice” and Hausa “Ina cin shinkafa”—to be mapped to distinct embeddings, severely hindering cross-lingual transfer for low-resource languages. To address this, we propose a parallel tokenizer framework: first training monolingual tokenizers independently, then aligning their vocabularies at the lexical level using bilingual dictionaries to enable semantically consistent cross-lingual subword sharing; finally constructing a unified, frequency-balanced shared semantic space. This work is the first to systematically reformulate vocabulary design in cross-lingual pretraining, enabling end-to-end Transformer pretraining. Pretrained on 13 low-resource languages, our model significantly outperforms mBERT and XLM-R on downstream tasks—including sentiment analysis and hate speech detection—demonstrating that vocabulary alignment is fundamental to cross-lingual generalization.

Addressing ineffective cross-lingual transfer in multilingual language modelsAligning vocabularies for consistent semantic representations across languagesImproving multilingual performance in low-resource language settings

This work addresses the inconsistency in reasoning exhibited by multilingual large language models when presented with semantically equivalent prompts in different languages, a problem often exacerbated by discrete tokenization. To mitigate this issue, the authors propose SOLAR, a novel approach that introduces soft tokens—probabilistic weighted combinations of lexical embeddings—as continuous semantic representations during supervised fine-tuning. By aligning multilingual soft token spaces through English as a pivot, SOLAR reduces reliance on specific vocabularies or writing systems while preserving cross-lingually shared reasoning structures. Experimental results demonstrate that SOLAR significantly outperforms baseline methods across four multilingual reasoning benchmarks, achieving gains of up to 17.7 percentage points, with particularly pronounced improvements for low-resource languages. The method also substantially enhances cross-lingual representational similarity.

cross-lingual reasoninglanguage-specific generationmultilingual large language models

Traditional Byte-Pair Encoding (BPE) tokenization introduces token redundancy in low-resource languages, degrading the performance of small-scale models. Method: This paper proposes a BPE configuration method integrating hyperparameter optimization and compressed sensing. It systematically searches key BPE hyperparameters—including vocabulary size and merge iterations—and jointly evaluates configurations using intrinsic metrics (e.g., token count) and extrinsic task performance (generation and classification). Contribution/Results: The study provides the first empirical evidence that BPE configuration significantly impacts multilingual modeling for low-resource languages. Experiments across diverse languages and model scales show that optimal configurations reduce token counts by 12.7% on average and improve downstream task accuracy by 1.8–3.4 percentage points for small models. These gains substantially enhance modeling efficiency and generalization capability in low-resource settings.

Compression-optimized tokenization benefits low-resource languagesImproved performance in multilingual NLP tasksOptimal BPE configuration reduces token count

A Text is Worth Several Tokens: Text Embedding from LLMs Secretly Aligns Well with The Key Tokens

Jun 25, 2024
ZN
Zhijie Nie
🏛️ Beihang University | Zhongguancun Laboratory

This work addresses the low efficiency of key information identification and semantic retrieval in text. We first discover that large language model (LLM) text embeddings naturally align with salient input tokens in the latent space—a phenomenon empirically validated across diverse model architectures, training paradigms, and embedding methods. Leveraging this insight, we propose a principal-component-guided alignment enhancement method that decomposes the embedding geometry and explicitly steers representations toward critical tokens. Building upon this, we design a lightweight sparse retrieval paradigm that retains only ~20% of token embeddings while achieving 80% of dense retrieval performance. Experiments across eight mainstream LLM embedders confirm the robustness of the alignment mechanism. Our findings provide an interpretable foundation for sparse retrieval and instruction-tuned embeddings, and advance the understanding of the intrinsic nature of semantic relevance.

Information ExtractionSemantic UnderstandingText Mining

Latest Papers

What's happening recently
View more

Subword Tokenization Strategies for Kurdish Word Embeddings

Nov 18, 2025
AS
Ali Salehi
🏛️ University at Buffalo

This study investigates how subword segmentation strategies affect morphological similarity modeling in low-resource Kurmanji Kurdish. Method: We propose a bootstrapped BiLSTM-CRF morphological segmenter trained on minimal human annotations and integrate it with Word2Vec to generate word embeddings. We systematically compare word-level, morpheme-level, and Byte-Pair Encoding (BPE) segmentation under a multi-dimensional evaluation framework assessing similarity preservation, clustering quality, and semantic structure fidelity. Contribution/Results: We identify a critical coverage bias in existing evaluations—BPE covers only 28.6% of test instances, leading to inflated performance estimates. In contrast, morpheme-level segmentation achieves 68.7% coverage and significantly outperforms BPE in semantic neighborhood quality and balanced representation of morphological complexity. These findings underscore the necessity of coverage-aware evaluation for low-resource NLP tasks, revealing that segmentation granularity and lexical coverage jointly determine embedding efficacy in morphologically rich, data-scarce languages.

Analyzing evaluation biases in tokenization method comparisonsDeveloping morphological segmenter with minimal manual annotationEvaluating tokenization strategies for Kurdish word embeddings

This study systematically investigates the intrinsic semantic and syntactic properties of mainstream word embedding methods—such as Word2Vec and GloVe—and their performance disparities across diverse natural language processing tasks. By establishing a unified evaluation framework that integrates publicly available pretrained embeddings with standard benchmark datasets, the work conducts empirical comparisons on canonical tasks including semantic similarity and analogical reasoning. The findings delineate the performance boundaries and optimal application scenarios for each embedding model, offering practitioners reliable guidance for model selection in real-world settings. Furthermore, the analysis deepens the understanding of the inherent limitations of static word representations, highlighting critical constraints in capturing contextual and compositional linguistic phenomena.

empirical investigationnatural language processingvector representations

How Good is BLI as an Alignment Measure: A Study in Word Embedding Paradigm

Nov 17, 2025
KW
Kasun Wickramasinghe
🏛️ University of Moratuwa

This study systematically evaluates bilingual lexicon induction (BLI) as a metric for assessing embedding space alignment quality, identifying its validity and limitations. Addressing diverse language pairs—including high- and low-resource settings and typologically distinct language families—the authors compare BLI performance across traditional linear alignment methods, multilingual pretrained models (e.g., mBERT, XLM-R), and novel compositional alignment strategies. They propose a stem-based BLI approach to improve matching accuracy for inflectional languages and introduce a more robust lexical pruning mechanism. Results show that compositional methods generally achieve the best BLI performance, while multilingual models exhibit pronounced advantages for low-resource languages. Crucially, standard BLI is found to be sensitive to vocabulary coverage bias and morphological noise, thus failing to reliably reflect true alignment quality. The work establishes a more rigorous methodological framework for evaluating embedding alignment, supported by extensive empirical analysis across linguistic settings.

Assessing alignment methods across high-resource and low-resource language pairsComparing performance of traditional, multilingual and combined alignment techniquesEvaluating BLI as a measure for embedding space alignment quality

This study addresses the unclear impact of tokenizer selection on multilingual model performance across languages, particularly low-resource ones. To investigate this, we conduct large-scale controlled experiments by training 123 models using 54 distinct tokenizers. We introduce a systematic evaluation framework incorporating the bits-per-byte metric, Spearman correlation analysis, and attribute-based predictive models. Our findings reveal how tokenizer influence varies with data scale, demonstrating that low-resource languages exhibit substantially greater sensitivity to tokenizer choice. Furthermore, we propose a candidate filtering strategy based on intrinsic tokenizer properties, enabling precise prediction of downstream model performance rankings through specific metrics. This work provides actionable insights for optimizing tokenizer selection in multilingual modeling, bridging the gap between intrinsic tokenizer characteristics and extrinsic task performance.

Bits-per-ByteLow-Resource LanguagesMultilingual Language Models

Hot Scholars

SU

Shubham Ugare

University of Illinois
Programming LanguagesMachine LearningCompilers
TS

Tarun Suresh

Undergraduate, University of Illinois Urbana-Champaign
Deep LearningMachine LearningReinforcement LearningProgramming Languages
PS

Philip S. Yu

Professor of Computer Science, University of Illinons at Chicago
Data miningDatabasePrivacy
SM

Sasa Misailovic

University of Illinois at Urbana–Champaign
Programming LanguagesCompilersApproximate ComputingProbabilistic Programming
XL

Xiaoze Liu

PhD Student, ECE at Purdue University
trustworthy ai