Multilingual GSM-Symbolic: What determines capability transfer across languages?
This study addresses the unclear mechanisms of cross-lingual capability transfer and the incomparability of evaluation data by constructing a multilingual symbolic mathematics dataset and proposing a joint estimation framework. Methodologically, it employs symbolic template generation to ensure data diversity, prevent overfitting, and systematically quantify key factors influencing capability transfer. The findings reveal that model scale and resource availability dominate cross-lingual transfer. Notably, the proposed framework explains 92% of the performance variance across languages and achieves a prediction error of merely six percentage points on unseen languages. Overall, this work provides a reliable theoretical foundation and an evaluation paradigm for understanding cross-lingual capability transfer in large language models.