🤖 AI Summary
This study addresses systematic biases in evaluating the grammatical competence of multilingual language models, which frequently arise from disparities in evaluation paradigms, post-training procedures, and resource availability. Leveraging the MultiBLiMP syntactic minimal pair benchmark, this work conducts a comparative analysis of six model families across four evaluation paradigms and introduces a cross-lingual native-language prompting technique. The findings reveal that post-training significantly degrades models' grammatical capabilities, with low-resource languages being disproportionately affected. Crucially, this grammatical knowledge is not erased but rather obscured, and can be effectively recovered through native-language prompting. Accordingly, this work proposes a language-aware, multi-paradigm evaluation protocol, offering a novel methodological framework for accurately measuring the true grammatical competence of multilingual models.
📝 Abstract
Claims about the grammatical competence of multilingual language models vary sharply with how competence is measured, yet the interaction between evaluation paradigm, post-training, and language resource availability has not been systematically examined. We evaluate base and post-trained models from six families on MultiBLiMP, a syntactic minimal-pair benchmark covering 101 languages, using four evaluation methods. We report three principal findings. First, post-training degrades grammatical competence, but the magnitude of this effect is reduced unevenly by model scale, while low-resource languages bear the highest cost. Second, post-trained models retain grammatical knowledge they cannot articulate through explicit prompting, yet this is measurable only in high-resource languages, because near-chance baselines in low-resource settings leave little knowledge to hide. Third, native-language prompting recovers otherwise hidden competence on low-resource languages, demonstrating that only high-resource languages can be probed directly from unprompted probabilities. We conclude that multilingual grammatical evaluation must adopt language-informed, multi-paradigm protocols to avoid systematically underestimating low-resource abilities.