π€ AI Summary
This study addresses the reliance of cross-lingual transfer performance evaluation on expensive multilingual pre-training, which lacks cost-effective predictive approaches. To overcome this limitation, this work proposes leveraging publicly available typological features as proxy signals to estimate transferability with zero computational overhead via a random forest model, thereby replacing hundreds of fine-tuning experiments. Through leave-one-out validation alongside script and language-family debiasing analyses, it demonstrates that typological signals can effectively reconstruct transfer matrices and decouple resource biases. Evaluated across 24 languages, the proposed approach achieves Ο=0.705 and RΒ²=0.49, significantly outperforming non-typological baselines. These findings establish typological features as an efficient and reliable tool for screening cross-lingual transfer candidates.
π Abstract
Cross-lingual transfer describes how knowledge in a source language benefits a target language. Measuring it quantitatively requires broad multilingual pre-training, as prior work has done with cross-lingual transfer matrices. We ask whether transfer is predictable from freely available typological features, and whether the prominence of high-resource source languages reflects typology or data quality and quantity. We show that typological databases contain cheap and dense signals about cross-lingual transfer. Our typology-only random forest on a 24-language prior-work transfer matrix scores leave-one-language-out $Ο{=}0.705$ and $R^2{=}0.49$, beating a non-typological control at $Ο{=}0.62$, which verifies the ability of typology-only predictions to reconstruct costly measured cross-lingual transfer. The signal survives leave-one-script-out and leave-one-family-out protocols, so script and family confounding do not explain the effect. By decomposing the transfer into a typology term and a resource-and-script bias term, we find the best-source ranking sensitive to this bias. In contrast, typology is not affected by this bias, which makes it a zero-compute screening tool that replaces hundreds of training runs with a model fit. Our code is available \href{https://github.com/dharmsen/typo-x-ling-transfer}{here}.