cross-lingual embedding comparison

Designs and implements representations and analysis pipelines that embed text or model outputs from multiple languages into a shared vector space, enabling computation of cross-language similarity metrics, clustering, and visualization. Also develops evaluation and correction methods to control for translation-induced biases and to compare languages or multilingual models robustly.

cross-lingualembeddingcomparison

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.46
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the reliance of multilingual large language models (mLLMs) on large-scale bilingual corpora and computationally intensive fine-tuning for cross-lingual alignment. We propose a fine-tuning-free, data-efficient intervention method that activates language-specific neurons in the embedding layer. By systematically identifying and manipulating these “language-expert” neurons, we analyze the geometric impact of such interventions on the embedding space—revealing, for the first time, that targeted activation can directionally enhance cross-lingual representation alignment. On cross-lingual retrieval, our method achieves up to a 2× improvement in top-1 accuracy. Crucially, we observe a strong correlation between embedding-space alignment metrics and downstream performance gains, confirming both the effectiveness and interpretability of the intervention mechanism. The approach substantially reduces dependence on bilingual supervision and training resources, establishing a lightweight, scalable paradigm for cross-lingual alignment.

Cross-lingual alignment in multilingual language modelsImproving cross-lingual retrieval task performanceModel interventions for embedding space manipulation

Words as Bridges: Exploring Computational Support for Cross-Disciplinary Translation Work

Mar 24, 2025
CB
Calvin Bao
🏛️ University of Maryland | MIT

Interdisciplinary literature exploration is often impeded by terminological barriers across domains. This paper conceptualizes disciplines as heterogeneous linguistic communities and, for the first time, adapts unsupervised cross-lingual word embedding alignment techniques to interdisciplinary concept alignment—preserving domain-specific terms as cognitive bridges rather than eliminating or oversimplifying them. Our method comprises: (1) domain-specific word vector training; (2) unsupervised alignment of embedding spaces across disciplines; and (3) construction and interactive design of a prototype concept-level cross-domain search engine. Evaluated in two case studies, our approach demonstrates effective semantic mapping, enabling concept-level—rather than lexical-level—cross-domain retrieval. Key contributions are: (1) establishing a novel paradigm for interdisciplinary concept alignment; (2) empirically validating domain terms as computationally tractable conceptual anchors; and (3) providing evidence-based insights for scholar-centered information-seeking interface design.

Aligning domain-specific word embeddings for cross-disciplinary conceptual explorationBridging jargon gaps between disciplines for better translationDeveloping computational tools to aid cross-domain information seeking

This study addresses the lack of reliable source language selection methods for cross-lingual transfer in low-resource African languages. Through a systematic evaluation of five embedding similarity metrics—cosine distance, P@1, CSLS, CKA, and others—across 816 cross-lingual transfer experiments spanning 12 African languages, three NLP tasks, and three Africa-centric multilingual models, the work demonstrates that cosine distance and retrieval-based metrics (P@1, CSLS) effectively predict transfer performance (Spearman’s ρ = 0.4–0.6), matching the predictive power of URIEL typological features. In contrast, CKA exhibits negligible predictive ability (ρ ≈ 0.1). The paper further presents the first direct comparison between embedding-based metrics and linguistic typology, uncovering a Simpson’s paradox when aggregating results across models, thereby underscoring the necessity of validating metric efficacy separately for each model.

African languagescross-lingual transferembedding similarity

Parallel Tokenizers: Rethinking Vocabulary Design for Cross-Lingual Transfer

Oct 07, 2025
MD
Muhammad Dehan Al Kautsar
🏛️ Mohamed bin Zayed University of Artificial Intelligence

Current multilingual tokenizers suffer from misaligned cross-lingual vocabularies, causing semantically equivalent words—e.g., English “I eat rice” and Hausa “Ina cin shinkafa”—to be mapped to distinct embeddings, severely hindering cross-lingual transfer for low-resource languages. To address this, we propose a parallel tokenizer framework: first training monolingual tokenizers independently, then aligning their vocabularies at the lexical level using bilingual dictionaries to enable semantically consistent cross-lingual subword sharing; finally constructing a unified, frequency-balanced shared semantic space. This work is the first to systematically reformulate vocabulary design in cross-lingual pretraining, enabling end-to-end Transformer pretraining. Pretrained on 13 low-resource languages, our model significantly outperforms mBERT and XLM-R on downstream tasks—including sentiment analysis and hate speech detection—demonstrating that vocabulary alignment is fundamental to cross-lingual generalization.

Addressing ineffective cross-lingual transfer in multilingual language modelsAligning vocabularies for consistent semantic representations across languagesImproving multilingual performance in low-resource language settings

Exploring Multilingual Probing in Large Language Models: A Cross-Language Analysis

Sep 22, 2024
DL
Daoyang Li
🏛️ University of Southern California | Rutgers University | Northwestern University | New Jersey Institute of Technology

Large language models (LLMs) exhibit weak and imbalanced structural knowledge representation—particularly for part-of-speech (POS) and dependency relations—across low-resource versus high-resource languages. Method: We propose a cross-lingual linear probing framework to systematically analyze the distribution and layer-wise evolution of syntactic knowledge in multilingual LLM representations. Our analysis spans high- and low-resource language groups, employing layer-wise accuracy tracking, cosine similarity measurements, and multilingual comparative experiments. Contribution/Results: We identify three systematic disparities for the first time: (1) significantly lower probing accuracy on low-resource languages; (2) diminished gains from deeper layers and flatter layer-wise accuracy trends; and (3) substantially reduced intra- and cross-group representation similarity. These findings reveal that structural knowledge transfer capability is critically resource-dependent—a core bottleneck in multilingual modeling. Our work provides empirical grounding and concrete directions for developing more equitable multilingual representation learning strategies.

Accuracy DisparityMultilingual Language ModelsResource-imbalanced Languages

Latest Papers

What's happening recently
View more

This study systematically investigates the intrinsic semantic and syntactic properties of mainstream word embedding methods—such as Word2Vec and GloVe—and their performance disparities across diverse natural language processing tasks. By establishing a unified evaluation framework that integrates publicly available pretrained embeddings with standard benchmark datasets, the work conducts empirical comparisons on canonical tasks including semantic similarity and analogical reasoning. The findings delineate the performance boundaries and optimal application scenarios for each embedding model, offering practitioners reliable guidance for model selection in real-world settings. Furthermore, the analysis deepens the understanding of the inherent limitations of static word representations, highlighting critical constraints in capturing contextual and compositional linguistic phenomena.

empirical investigationnatural language processingvector representations

This study addresses the lack of systematic evaluation of performance stability across tasks and languages in existing large-scale multilingual text embedding models, noting that benchmark conclusions are often sensitive to dataset composition and aggregation methodologies. To this end, the authors propose a meta-research framework based on multi-criteria decision-making ranking, enabling robust cross-task and cross-lingual analysis of models covering approximately 230 languages on the MTEB benchmark. The framework introduces two novel metrics—“dataset composition robustness” and “ranking scheme robustness”—to facilitate systematic sensitivity assessment of benchmark findings. Results reveal that while large models generally exhibit stable performance across most tasks, notable exceptions arise in retrieval tasks; moreover, only a handful of models consistently outperform others across diverse tasks, ranking strategies, and dataset subsets.

benchmark robustnesscross-task generalizationevaluation methodology

This work addresses the barriers to equitable development of high-quality text embeddings—namely high costs, limited language coverage, and model opacity—by introducing a three-dimensional Matryoshka Learning framework (3D-ML). This novel approach integrates Matryoshka learning across representation, layer, and embedding dimensions, establishing Matryoshka Embedding Learning (MEL) to achieve breakthroughs in parameter efficiency, inference flexibility, and storage compression. Leveraging this framework, the authors develop a large-scale multilingual embedding model covering over 100 languages and release the model, training data, and code publicly. Comprehensive evaluation across 430 tasks demonstrates state-of-the-art performance, setting new records on 9 out of 17 MTEB benchmarks, with particularly strong results on low-resource languages.

AI equitycomputational costlinguistic inclusivity

Hot Scholars

RT

Reut Tsarfaty

Bar-Ilan University
Natural Language ProcessingComputational LinguisticsArtificial Inteligence
YL

Yuanchao Li

University of Edinburgh
speech technologiesspoken language processingaffective computingdigital health
PL

Peerat Limkonchotiwat

Research Fellow, AI Singapore, National University of Singapore
Evaluation and BenchmarkRepresentation LearningLarge Language ModelMultilingual Learning
AF

Alham Fikri Aji

MBZUAI, Monash Indonesia
MultilingualityLow-resource NLPLanguage ModelingMachine Translation
YO

Yohei Oseki

University of Tokyo
Computational LinguisticsCognitive Science