Language Distances are Practical for Equitable Cross-Lingual Transfer

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of source language selection and fairness in low-resource scenarios within cross-lingual transfer. It presents the first systematic evaluation of language distance-based source language ranking paradigms from a fairness perspective. By integrating multilingual models, training-free composite distance metrics, and supervised learning-to-rank algorithms, this work thoroughly investigates the trade-off between task performance and resource inequality. The findings validate the effectiveness of language distance as a criterion for source language selection and demonstrate that trained rankers offer significant advantages in mitigating inequality. Accordingly, we recommend prioritizing trained rankers when evaluation data are available, and resorting to composite distance metrics otherwise, thereby achieving an optimal balance between performance and fairness in cross-lingual transfer.
📝 Abstract
Cross-lingual transfer is strongly conditional on how the source language is chosen, but it is impractical to determine the best candidate source for every target language, especially for low-resource target languages. Language distances are widely used to rank candidate sources due to their correlation with transfer efficacy and applicability in resource-sparse settings. However, the reliability of distance-based rankers across tasks and resource levels remains underexplored. We therefore present the first equity-focused evaluation of paradigms for ranking source languages, studying resource-level inequality and task inequality across ten cross-lingual tasks and two multilingual models. While both inequalities are most pronounced for individual language distances and an English-always baseline, they are substantially reduced by training-free composite distances, and nearly eliminated by trained rankers. We further demonstrate the reliability of rankers using language distances compared to rankers using language model internals. Overall, we find that language distances provide a practical basis for equitable and performant transfer language selection. We recommend using trained rankers when task-specific transfer evaluations are available, and composite distances otherwise.
Problem

Research questions and friction points this paper is trying to address.

cross-lingual transfer
language distances
source language selection
equity
low-resource languages
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-lingual Transfer
Language Distances
Equity-focused Evaluation
Composite Distances
Trained Rankers
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
York Hay Ng
University of Toronto, Canada
R
Razan Ahsan Rifandi
University of Toronto, Canada
A
Aditya Khan
University of Toronto, Canada
En-Shiun Annie Lee
En-Shiun Annie Lee
Ontario Tech University, and University of Toronto (Status-Only)
Natural Language ProcessingData MiningPattern Analysis