What Does Post-Training Change in Multilingual Reasoning?

๐Ÿ“… 2026-09-29
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the significant performance gap of open-source reasoning models in multilingual scenarios, where non-English accuracy lags substantially behind English. We systematically audit the multilingual reasoning bottlenecks of Qwen3, revealing that post-training constitutes the dominant limiting factor. To overcome this, we propose a reinforcement learning framework incorporating multilingual supervised fine-tuning and language-term rewards, which effectively mitigates non-termination loops and language fallback issues. This work establishes a constructive post-training pathway that transfers English-centric capabilities to multilingual reasoning, significantly improving both answer completeness and accuracy in non-English settings.
๐Ÿ“ Abstract
Open-source reasoning models provide unequal access to reasoning capability across languages. When a model can solve a problem but cannot deliver a complete solution in the user's language, language becomes an access barrier rather than merely a source of performance variation. We audit Qwen3 checkpoints on competition-mathematics tasks in eleven languages. Across the ten non-English languages, only 15.4-17.9% of problems receive a correct, terminating solution with visible reasoning in the requested language in any of 16 samples, compared with 92.9% in English. To identify the source of this disparity, we evaluate thirteen endpoints from one model family, spanning released checkpoints, multilingual supervised fine-tuning (SFT) at two scales, controlled SFT ablations, and three reinforcement-learning (RL) reward formulations. We jointly track correctness, language adherence, termination, and delivery efficiency. The dominant bottleneck shifts across post-training stages. Released models often reason in English. Multilingual SFT restores target-language reasoning, but accuracy declines across multilingual, English-only, and single-language SFT runs, showing that this cost is not specific to multilingual mixing; non-English reasoning traces additionally become prone to non-terminating loops. RL restores termination in both arms at no cost in accuracy, but only the arm whose reward includes a language term delivers: rewarding correctness alone returns the model to English. Together, these stages establish a constructive post-training path from English-pivoted capability to multilingual reasoning that is reliably delivered.
Problem

Research questions and friction points this paper is trying to address.

multilingual reasoning
post-training
language disparity
open-source reasoning models
reinforcement learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multilingual Reasoning
Post-Training
Supervised Fine-Tuning
Reinforcement Learning
Reward Formulation
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.