Score
Designs, implements, and evaluates models, representation spaces, and adaptation procedures that enable knowledge learned in one language to be reused in others; this includes training and aligning multilingual or cross-lingual embeddings, fine-tuning or adapting models across languages, building or integrating translation-based transfer components, and measuring cross-lingual similarity and robustness. As a practitioner you build the transfer pipelines and evaluation protocols (metrics, target-language tests) needed to assess and improve cross-lingual generalization.
This study addresses the limited generalization of multilingual neural machine translation (NMT) for low-resource languages, which often stems from the scarcity of parallel data. The authors systematically investigate how linguistic similarity, data composition, and training strategies influence cross-lingual knowledge transfer. They propose a novel approach that integrates retrieval-augmented mechanisms with auxiliary supervision signals and further analyze performance trade-offs during fine-tuning. Experimental results demonstrate that the proposed method substantially improves translation quality for low-resource languages, enhances model generalization, and reduces out-of-domain generation. These findings offer an effective pathway toward building more robust and inclusive multilingual natural language processing systems.
Existing evaluation methods struggle to disentangle overall performance gains in source languages from genuine cross-lingual transfer capabilities in multilingual models. To address this limitation, this work proposes the Hardness-Adjusted Transfer (HAT) score, which isolates source-language performance to more accurately quantify transfer effectiveness from high-resource to low-resource languages. Leveraging HAT, we conduct a large-scale empirical analysis across 20 language models and three major multilingual benchmarks, revealing—for the first time—that small models retain meaningful transfer capacity, that scaling model size yields diminishing returns in transfer gains, and that overall cross-lingual transfer capability has steadily improved over time.
This study investigates the dynamic mechanisms of cross-lingual transfer (CLT) in large language models (35B parameters) under realistic post-training scenarios, focusing on multilingual generation across summarization, instruction following, and mathematical reasoning. Method: We conduct systematic analysis under both single-task and multi-task instruction tuning regimes, employing controlled multilingual instruction data, cross-lingual performance attribution, and large-scale fine-tuning evaluation across Qwen and LLaMA model families. Contribution/Results: We首次 uncover that CLT exhibits nonlinear dependence on data mixing ratios, task complexity, and training paradigm combinations. We propose a reproducible efficacy criterion for CLT and identify optimal data proportioning and task-scheduling strategies that significantly enhance low-resource language performance—achieving a 27% absolute zero-shot cross-lingual accuracy gain in mathematical reasoning.
Large language models (LLMs) exhibit degraded performance on low-resource programming languages (e.g., COBOL, Rust, Swift) due to insufficient training data. Method: This paper systematically investigates cross-lingual transfer learning, introducing the first large-scale empirical framework for characterizing transfer patterns across programming languages—evaluated across 11–41 languages and 1,808 task-language combinations on code completion, translation, and repair. Contribution/Results: We empirically identify Kotlin and JavaScript as optimal source languages; uncover task-specific, heterogeneous dependencies on source-language features—challenging natural-language transfer paradigms; and develop both a principled source-language selection guide and a feature-based prediction model. Our approach significantly improves performance on low-resource languages across diverse coding tasks, establishing a scalable methodology for legacy system modernization and AI support for emerging programming languages.
This work addresses catastrophic forgetting in cross-lingual transfer—specifically, the degradation of source-language knowledge during target-language fine-tuning. We propose Cross-Lingual Validation (CLV), a novel paradigm that, for the first time within a unified framework, quantifies forgetting magnitude across multilingual models. Systematically comparing full-parameter fine-tuning versus adapter-based tuning, and intermediate-task training (IT) versus CLV, we analyze their trade-offs in preserving source-language (English) performance versus optimizing target-language accuracy. Experiments employ large language models on multilingual hate speech detection and product review classification datasets under zero-shot and full-shot settings. Results show CLV significantly outperforms IT in retaining source-language knowledge, reducing catastrophic forgetting by 23.6% average F1—challenging the prevailing assumption of IT’s superiority. Although IT yields marginally higher target-language performance, CLV achieves superior cross-lingual stability and transferability.
Pretraining contamination undermines the evaluation of cross-lingual knowledge transfer in multilingual large language models (LLMs). Method: We propose a time-sensitive, automated evaluation framework that mines entity facts relative to temporal knowledge cutoff points, aligns cross-lingual documents, and automatically generates questions—yielding a rigorously validated multilingual benchmark with strict knowledge-cutoff enforcement to isolate true cross-lingual transfer from pretraining exposure. Contribution/Results: Our framework enables the first precise measurement of genuine cross-lingual knowledge transfer, uncovering migration asymmetry induced by linguistic distance and diminishing marginal returns with increasing model scale. Evaluated across five languages and multiple state-of-the-art models, it establishes a reproducible, contamination-resistant benchmark for multilingual knowledge transfer assessment.
In multilingual domain adaptation (ML-DA), the intra-lingual acquisition mechanisms of domain knowledge and cross-lingual transfer pathways remain poorly understood, hindering performance on low-resource languages. Method: Focusing on the English–Japanese bilingual biomedical domain, this work systematically investigates knowledge acquisition dynamics in a 13B-parameter large language model. We propose AdaXEval—a structured, bilingual multiple-choice QA evaluation framework built on domain-specific parallel corpora—to enable fine-grained, continuous tracking of knowledge learning. Experiments employ continual training with multi-formulation data strategies. Contribution/Results: Despite high-quality bilingual data, cross-lingual knowledge transfer exhibits pronounced asymmetry. AdaXEval effectively uncovers transfer bottlenecks and intra-lingual knowledge consolidation patterns. All code and datasets are publicly released.
This work addresses the challenge of disentangling the mechanisms underlying cross-lingual generalization in language models, which is confounded in natural corpora by intertwined factors such as lexical overlap, morphological variation, and data imbalance. To isolate these variables, the authors propose an in vitro experimental framework that procedurally generates two synthetic languages sharing identical ontologies and syntactic structures but differing in surface forms. By systematically controlling lexical distance, minority-language proportion, tokenization strategy, and vocabulary size across 70 controlled experiments, they find that transfer performance hinges not on lexical similarity but on whether the tokenizer preserves reusable cross-lingual subword structures. A smaller vocabulary enhances decomposability of words, thereby improving masked language modeling transfer. Moreover, cross-lingual transfer exhibits a staged pattern, prioritizing grammar and typology over lexical alignment.
Prior scaling law studies are predominantly English-centric, neglecting multilingual settings. Method: This work systematically investigates scaling laws for multilingual models spanning 10M–8B parameters across 400+ languages, uncovering cross-lingual transfer mechanisms and the “multilinguality curse.” We propose Adaptive Transfer Scaling (ATLAS), which constructs a language-pair transfer matrix and identifies, for the first time, computational inflection points for zero-shot pretraining and fine-tuning—enabling language-agnostic optimal scaling. Contribution/Results: Based on 774 large-scale experiments, combined with regression analysis, cross-lingual performance prediction, and empirical transfer modeling, ATLAS improves R² over existing scaling laws by >0.3. It quantifies reciprocity across 1,444 language pairs and establishes a scalable multilingual training paradigm and resource-allocation principle for non-English-dominant scenarios.
This study addresses the lack of reliable source language selection methods for cross-lingual transfer in low-resource African languages. Through a systematic evaluation of five embedding similarity metrics—cosine distance, P@1, CSLS, CKA, and others—across 816 cross-lingual transfer experiments spanning 12 African languages, three NLP tasks, and three Africa-centric multilingual models, the work demonstrates that cosine distance and retrieval-based metrics (P@1, CSLS) effectively predict transfer performance (Spearman’s ρ = 0.4–0.6), matching the predictive power of URIEL typological features. In contrast, CKA exhibits negligible predictive ability (ρ ≈ 0.1). The paper further presents the first direct comparison between embedding-based metrics and linguistic typology, uncovering a Simpson’s paradox when aggregating results across models, thereby underscoring the necessity of validating metric efficacy separately for each model.