perform cross-lingual transfer

Designs, implements, and evaluates models, representation spaces, and adaptation procedures that enable knowledge learned in one language to be reused in others; this includes training and aligning multilingual or cross-lingual embeddings, fine-tuning or adapting models across languages, building or integrating translation-based transfer components, and measuring cross-lingual similarity and robustness. As a practitioner you build the transfer pipelines and evaluation protocols (metrics, target-language tests) needed to assess and improve cross-lingual generalization.

performcross-lingualtransfer

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.73
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$190K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the limited generalization of multilingual neural machine translation (NMT) for low-resource languages, which often stems from the scarcity of parallel data. The authors systematically investigate how linguistic similarity, data composition, and training strategies influence cross-lingual knowledge transfer. They propose a novel approach that integrates retrieval-augmented mechanisms with auxiliary supervision signals and further analyze performance trade-offs during fine-tuning. Experimental results demonstrate that the proposed method substantially improves translation quality for low-resource languages, enhances model generalization, and reduces out-of-domain generation. These findings offer an effective pathway toward building more robust and inclusive multilingual natural language processing systems.

cross-lingual knowledge transferlow-resource languagesmultilingual machine translation

Existing evaluation methods struggle to disentangle overall performance gains in source languages from genuine cross-lingual transfer capabilities in multilingual models. To address this limitation, this work proposes the Hardness-Adjusted Transfer (HAT) score, which isolates source-language performance to more accurately quantify transfer effectiveness from high-resource to low-resource languages. Leveraging HAT, we conduct a large-scale empirical analysis across 20 language models and three major multilingual benchmarks, revealing—for the first time—that small models retain meaningful transfer capacity, that scaling model size yields diminishing returns in transfer gains, and that overall cross-lingual transfer capability has steadily improved over time.

cross-lingual transferevaluation metriclanguage representation

A Post-trainer's Guide to Multilingual Training Data: Uncovering Cross-lingual Transfer Dynamics

Apr 23, 2025
LS
Luísa Shimabucoro
🏛️ Cohere For AI | University of São Paulo | Meta

This study investigates the dynamic mechanisms of cross-lingual transfer (CLT) in large language models (35B parameters) under realistic post-training scenarios, focusing on multilingual generation across summarization, instruction following, and mathematical reasoning. Method: We conduct systematic analysis under both single-task and multi-task instruction tuning regimes, employing controlled multilingual instruction data, cross-lingual performance attribution, and large-scale fine-tuning evaluation across Qwen and LLaMA model families. Contribution/Results: We首次 uncover that CLT exhibits nonlinear dependence on data mixing ratios, task complexity, and training paradigm combinations. We propose a reproducible efficacy criterion for CLT and identify optimal data proportioning and task-scheduling strategies that significantly enhance low-resource language performance—achieving a 27% absolute zero-shot cross-lingual accuracy gain in mathematical reasoning.

Analyzing performance across tasks and model sizesIdentifying conditions for effective cross-lingual transferUnderstanding cross-lingual transfer dynamics in multilingual training

Learning Transfers over Several Programming Languages

Oct 25, 2023
RB
Razan Baltaji
🏛️ University of Illinois | IBM Research

Large language models (LLMs) exhibit degraded performance on low-resource programming languages (e.g., COBOL, Rust, Swift) due to insufficient training data. Method: This paper systematically investigates cross-lingual transfer learning, introducing the first large-scale empirical framework for characterizing transfer patterns across programming languages—evaluated across 11–41 languages and 1,808 task-language combinations on code completion, translation, and repair. Contribution/Results: We empirically identify Kotlin and JavaScript as optimal source languages; uncover task-specific, heterogeneous dependencies on source-language features—challenging natural-language transfer paradigms; and develop both a principled source-language selection guide and a feature-based prediction model. Our approach significantly improves performance on low-resource languages across diverse coding tasks, establishing a scalable methodology for legacy system modernization and AI support for emerging programming languages.

Enhancing LLM performance for low-resource programming languagesInvestigating transfer learning across 10-41 programming languagesPredicting optimal source languages for cross-lingual transfer

This work addresses catastrophic forgetting in cross-lingual transfer—specifically, the degradation of source-language knowledge during target-language fine-tuning. We propose Cross-Lingual Validation (CLV), a novel paradigm that, for the first time within a unified framework, quantifies forgetting magnitude across multilingual models. Systematically comparing full-parameter fine-tuning versus adapter-based tuning, and intermediate-task training (IT) versus CLV, we analyze their trade-offs in preserving source-language (English) performance versus optimizing target-language accuracy. Experiments employ large language models on multilingual hate speech detection and product review classification datasets under zero-shot and full-shot settings. Results show CLV significantly outperforms IT in retaining source-language knowledge, reducing catastrophic forgetting by 23.6% average F1—challenging the prevailing assumption of IT’s superiority. Although IT yields marginally higher target-language performance, CLV achieves superior cross-lingual stability and transferability.

Compare fine-tuning strategies for language modelsEvaluate knowledge retention across multiple languagesMeasure catastrophic forgetting in cross-lingual transfer

Latest Papers

What's happening recently
View more

LiveCLKTBench: Towards Reliable Evaluation of Cross-Lingual Knowledge Transfer in Multilingual LLMs

Nov 03, 2025
PG
Pei-Fu Guo
🏛️ National Taiwan University | University of California, Los Angeles | Academia Sinica

Pretraining contamination undermines the evaluation of cross-lingual knowledge transfer in multilingual large language models (LLMs). Method: We propose a time-sensitive, automated evaluation framework that mines entity facts relative to temporal knowledge cutoff points, aligns cross-lingual documents, and automatically generates questions—yielding a rigorously validated multilingual benchmark with strict knowledge-cutoff enforcement to isolate true cross-lingual transfer from pretraining exposure. Contribution/Results: Our framework enables the first precise measurement of genuine cross-lingual knowledge transfer, uncovering migration asymmetry induced by linguistic distance and diminishing marginal returns with increasing model scale. Evaluated across five languages and multiple state-of-the-art models, it establishes a reproducible, contamination-resistant benchmark for multilingual knowledge transfer assessment.

Analyzing how linguistic distance and model scale affect multilingual transferEvaluating transferability across languages using time-sensitive factual questionsIsolating genuine cross-lingual knowledge transfer from prior exposure in LLMs

Tracing Multilingual Knowledge Acquisition Dynamics in Domain Adaptation: A Case Study of English-Japanese Biomedical Adaptation

Oct 13, 2025
XZ
Xin Zhao
🏛️ The University of Tokyo | Institute of Industrial Science | National Institute of Informatics

In multilingual domain adaptation (ML-DA), the intra-lingual acquisition mechanisms of domain knowledge and cross-lingual transfer pathways remain poorly understood, hindering performance on low-resource languages. Method: Focusing on the English–Japanese bilingual biomedical domain, this work systematically investigates knowledge acquisition dynamics in a 13B-parameter large language model. We propose AdaXEval—a structured, bilingual multiple-choice QA evaluation framework built on domain-specific parallel corpora—to enable fine-grained, continuous tracking of knowledge learning. Experiments employ continual training with multi-formulation data strategies. Contribution/Results: Despite high-quality bilingual data, cross-lingual knowledge transfer exhibits pronounced asymmetry. AdaXEval effectively uncovers transfer bottlenecks and intra-lingual knowledge consolidation patterns. All code and datasets are publicly released.

Addressing suboptimal performance in low-resource multilingual domain adaptationExamining multilingual knowledge acquisition dynamics in domain adaptationInvestigating cross-lingual knowledge transfer mechanisms in LLMs

This work addresses the challenge of disentangling the mechanisms underlying cross-lingual generalization in language models, which is confounded in natural corpora by intertwined factors such as lexical overlap, morphological variation, and data imbalance. To isolate these variables, the authors propose an in vitro experimental framework that procedurally generates two synthetic languages sharing identical ontologies and syntactic structures but differing in surface forms. By systematically controlling lexical distance, minority-language proportion, tokenization strategy, and vocabulary size across 70 controlled experiments, they find that transfer performance hinges not on lexical similarity but on whether the tokenizer preserves reusable cross-lingual subword structures. A smaller vocabulary enhances decomposability of words, thereby improving masked language modeling transfer. Moreover, cross-lingual transfer exhibits a staged pattern, prioritizing grammar and typology over lexical alignment.

cross-lingual generalizationdata imbalancelanguage models

ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality

Oct 24, 2025
SL
Shayne Longpre
🏛️ MIT | University of Washington | Stanford University | Google Cloud AI | Google DeepMind

Prior scaling law studies are predominantly English-centric, neglecting multilingual settings. Method: This work systematically investigates scaling laws for multilingual models spanning 10M–8B parameters across 400+ languages, uncovering cross-lingual transfer mechanisms and the “multilinguality curse.” We propose Adaptive Transfer Scaling (ATLAS), which constructs a language-pair transfer matrix and identifies, for the first time, computational inflection points for zero-shot pretraining and fine-tuning—enabling language-agnostic optimal scaling. Contribution/Results: Based on 774 large-scale experiments, combined with regression analysis, cross-lingual performance prediction, and empirical transfer modeling, ATLAS improves R² over existing scaling laws by >0.3. It quantifies reciprocity across 1,444 language pairs and establishes a scalable multilingual training paradigm and resource-allocation principle for non-English-dominant scenarios.

Developing scaling laws for multilingual pretraining beyond English dominanceOptimizing model scaling strategies to overcome multilinguality performance limitationsQuantifying cross-lingual transfer benefits across 1444 language pairs

This study addresses the lack of reliable source language selection methods for cross-lingual transfer in low-resource African languages. Through a systematic evaluation of five embedding similarity metrics—cosine distance, P@1, CSLS, CKA, and others—across 816 cross-lingual transfer experiments spanning 12 African languages, three NLP tasks, and three Africa-centric multilingual models, the work demonstrates that cosine distance and retrieval-based metrics (P@1, CSLS) effectively predict transfer performance (Spearman’s ρ = 0.4–0.6), matching the predictive power of URIEL typological features. In contrast, CKA exhibits negligible predictive ability (ρ ≈ 0.1). The paper further presents the first direct comparison between embedding-based metrics and linguistic typology, uncovering a Simpson’s paradox when aggregating results across models, thereby underscoring the necessity of validating metric efficacy separately for each model.

African languagescross-lingual transferembedding similarity

Hot Scholars

AF

Alham Fikri Aji

MBZUAI, Monash Indonesia
MultilingualityLow-resource NLPLanguage ModelingMachine Translation
DI

David Ifeoluwa Adelani

McGill University and Mila - Quebec AI Institute and Canada CIFAR AI Chair
Natural language processingMultilingualityMultilingual NLPAfricaNLP
GS

Grigori Sidorov

Professor of Computational Linguistics, Instituto Politécnico Nacional (IPN), Mexico
Computational LinguisticsNatural Language ProcessingArtificial IntelligenceMachine Learning
IA

Idris Abdulmumin

Postdoctoral Fellow, DSFSI, University of Pretoria
Machine TranslationNeural Machine TranslationNatural Language ProcessingInternet Technology
SH

Shamsuddeen Hassan Muhammad

Bayero University, Kano, & Google DeepMind Academic Fellow at Imperial College London
Natural Language ProcessingSentiment AnalysisAfricaNLPLow-resource NLP