Score
Designs and implements representations and analysis pipelines that embed text or model outputs from multiple languages into a shared vector space, enabling computation of cross-language similarity metrics, clustering, and visualization. Also develops evaluation and correction methods to control for translation-induced biases and to compare languages or multilingual models robustly.
This work addresses the reliance of multilingual large language models (mLLMs) on large-scale bilingual corpora and computationally intensive fine-tuning for cross-lingual alignment. We propose a fine-tuning-free, data-efficient intervention method that activates language-specific neurons in the embedding layer. By systematically identifying and manipulating these “language-expert” neurons, we analyze the geometric impact of such interventions on the embedding space—revealing, for the first time, that targeted activation can directionally enhance cross-lingual representation alignment. On cross-lingual retrieval, our method achieves up to a 2× improvement in top-1 accuracy. Crucially, we observe a strong correlation between embedding-space alignment metrics and downstream performance gains, confirming both the effectiveness and interpretability of the intervention mechanism. The approach substantially reduces dependence on bilingual supervision and training resources, establishing a lightweight, scalable paradigm for cross-lingual alignment.
Interdisciplinary literature exploration is often impeded by terminological barriers across domains. This paper conceptualizes disciplines as heterogeneous linguistic communities and, for the first time, adapts unsupervised cross-lingual word embedding alignment techniques to interdisciplinary concept alignment—preserving domain-specific terms as cognitive bridges rather than eliminating or oversimplifying them. Our method comprises: (1) domain-specific word vector training; (2) unsupervised alignment of embedding spaces across disciplines; and (3) construction and interactive design of a prototype concept-level cross-domain search engine. Evaluated in two case studies, our approach demonstrates effective semantic mapping, enabling concept-level—rather than lexical-level—cross-domain retrieval. Key contributions are: (1) establishing a novel paradigm for interdisciplinary concept alignment; (2) empirically validating domain terms as computationally tractable conceptual anchors; and (3) providing evidence-based insights for scholar-centered information-seeking interface design.
This study addresses the lack of reliable source language selection methods for cross-lingual transfer in low-resource African languages. Through a systematic evaluation of five embedding similarity metrics—cosine distance, P@1, CSLS, CKA, and others—across 816 cross-lingual transfer experiments spanning 12 African languages, three NLP tasks, and three Africa-centric multilingual models, the work demonstrates that cosine distance and retrieval-based metrics (P@1, CSLS) effectively predict transfer performance (Spearman’s ρ = 0.4–0.6), matching the predictive power of URIEL typological features. In contrast, CKA exhibits negligible predictive ability (ρ ≈ 0.1). The paper further presents the first direct comparison between embedding-based metrics and linguistic typology, uncovering a Simpson’s paradox when aggregating results across models, thereby underscoring the necessity of validating metric efficacy separately for each model.
Current multilingual tokenizers suffer from misaligned cross-lingual vocabularies, causing semantically equivalent words—e.g., English “I eat rice” and Hausa “Ina cin shinkafa”—to be mapped to distinct embeddings, severely hindering cross-lingual transfer for low-resource languages. To address this, we propose a parallel tokenizer framework: first training monolingual tokenizers independently, then aligning their vocabularies at the lexical level using bilingual dictionaries to enable semantically consistent cross-lingual subword sharing; finally constructing a unified, frequency-balanced shared semantic space. This work is the first to systematically reformulate vocabulary design in cross-lingual pretraining, enabling end-to-end Transformer pretraining. Pretrained on 13 low-resource languages, our model significantly outperforms mBERT and XLM-R on downstream tasks—including sentiment analysis and hate speech detection—demonstrating that vocabulary alignment is fundamental to cross-lingual generalization.
Large language models (LLMs) exhibit weak and imbalanced structural knowledge representation—particularly for part-of-speech (POS) and dependency relations—across low-resource versus high-resource languages. Method: We propose a cross-lingual linear probing framework to systematically analyze the distribution and layer-wise evolution of syntactic knowledge in multilingual LLM representations. Our analysis spans high- and low-resource language groups, employing layer-wise accuracy tracking, cosine similarity measurements, and multilingual comparative experiments. Contribution/Results: We identify three systematic disparities for the first time: (1) significantly lower probing accuracy on low-resource languages; (2) diminished gains from deeper layers and flatter layer-wise accuracy trends; and (3) substantially reduced intra- and cross-group representation similarity. These findings reveal that structural knowledge transfer capability is critically resource-dependent—a core bottleneck in multilingual modeling. Our work provides empirical grounding and concrete directions for developing more equitable multilingual representation learning strategies.
This study systematically investigates the intrinsic semantic and syntactic properties of mainstream word embedding methods—such as Word2Vec and GloVe—and their performance disparities across diverse natural language processing tasks. By establishing a unified evaluation framework that integrates publicly available pretrained embeddings with standard benchmark datasets, the work conducts empirical comparisons on canonical tasks including semantic similarity and analogical reasoning. The findings delineate the performance boundaries and optimal application scenarios for each embedding model, offering practitioners reliable guidance for model selection in real-world settings. Furthermore, the analysis deepens the understanding of the inherent limitations of static word representations, highlighting critical constraints in capturing contextual and compositional linguistic phenomena.
This study addresses the lack of systematic evaluation of performance stability across tasks and languages in existing large-scale multilingual text embedding models, noting that benchmark conclusions are often sensitive to dataset composition and aggregation methodologies. To this end, the authors propose a meta-research framework based on multi-criteria decision-making ranking, enabling robust cross-task and cross-lingual analysis of models covering approximately 230 languages on the MTEB benchmark. The framework introduces two novel metrics—“dataset composition robustness” and “ranking scheme robustness”—to facilitate systematic sensitivity assessment of benchmark findings. Results reveal that while large models generally exhibit stable performance across most tasks, notable exceptions arise in retrieval tasks; moreover, only a handful of models consistently outperform others across diverse tasks, ranking strategies, and dataset subsets.
This work addresses the barriers to equitable development of high-quality text embeddings—namely high costs, limited language coverage, and model opacity—by introducing a three-dimensional Matryoshka Learning framework (3D-ML). This novel approach integrates Matryoshka learning across representation, layer, and embedding dimensions, establishing Matryoshka Embedding Learning (MEL) to achieve breakthroughs in parameter efficiency, inference flexibility, and storage compression. Leveraging this framework, the authors develop a large-scale multilingual embedding model covering over 100 languages and release the model, training data, and code publicly. Comprehensive evaluation across 430 tasks demonstrates state-of-the-art performance, setting new records on 9 out of 17 MTEB benchmarks, with particularly strong results on low-resource languages.