🤖 AI Summary
This study investigates the relative contributions of linguistic relatedness and task alignment in cross-lingual transfer, challenging the prevailing assumption that performance gains primarily stem from language similarity. By fine-tuning large language models—including both dense and mixture-of-experts architectures—on Arabic and evaluating them zero-shot on reading comprehension tasks across Semitic languages and typologically distant controls, the authors conduct chain-of-thought ablation experiments to demonstrate, for the first time empirically, that performance improvements on related languages are driven predominantly by task alignment rather than phylogenetic proximity. The findings reveal that weaker baseline models exhibit substantial gains across all languages, whereas stronger baselines show limited improvement; furthermore, the benefits conferred by chain-of-thought reasoning at inference align closely with those from fine-tuning, reinforcing the central role of task alignment in effective cross-lingual transfer.
📝 Abstract
We study cross-lingual transfer by fine-tuning seven large language models (4B--671B parameters) on Arabic and evaluating zero-shot reading comprehension on Semitic languages and non-Semitic controls. Across dense and Mixture-of-Experts architectures, we find no evidence of Semitic-specific transfer: models with weak baselines improve dramatically across all languages, while strong-baseline models show only marginal gains regardless of language family. A chain-of-thought ablation reinforces this finding -- the same models that benefit most from fine-tuning benefit equally from inference-time reasoning, suggesting both mechanisms address task-format alignment rather than cross-lingual knowledge transfer.