π€ AI Summary
This study addresses the limitation of the single pass@1 metric in distinguishing whether mathematical reasoning gains in large language models stem from novel solution discovery, sampling compression, or memorization. To this end, it proposes a dual-state diagnostic framework based on large-K saturation, combined with cross-surface probing via multilingual and isomorphic prompts, to comparatively analyze how distillation and GRPO post-training pathways specifically influence reasoning mechanisms. The findings reveal an essential dichotomy: easy problems rely on compression efficiency, whereas hard problems depend on extending capability ceilings. Furthermore, the results demonstrate that sufficiently off-policy distillation can elevate performance ceilings on hard problems, surpassing DeepSeek-Math, while also indicating that although English-dominant distillation improves non-English reasoning, inherent linguistic gaps persist.
π Abstract
Post-training is central to mathematical reasoning in modern large language models (LLMs), but endpoint pass@1 alone underidentifies what has changed. Gains may reflect newly reachable solutions, cheaper sampling of latent solutions, surface robustness, or memorisation. We compare three post-training paths under a common diagnostic readout: our sufficiently trained off-policy distillation trajectories, released Qwen3 off-policy-plus-on-policy distillation endpoints, and a released DeepSeek-Math endpoint trained with Group Relative Policy Optimisation (GRPO). Our probe uses cross-surface pass@K over verbatim prompts, paraphrases, numerical isomorphisms, and translations, plus consistency, distribution-shape, and verified supervised-fine-tuning (SFT) membership analyses. We find two regimes. On easier AMC problems, large-K ceilings are near saturation, so post-training mainly compresses sample cost. On harder AIME problems, post-training expands the large-K ceiling over the base model: sufficient off-policy distillation already raises this ceiling, Qwen3 released endpoints raise it further, and DeepSeek-Math GRPO does not dominate sufficient off-policy distillation at large K. English-dominant distillation improves non-English reasoning but preserves language-tier gaps. A controlled-overfit audit finds limited sensitivity in current SFT-membership probes. Compression is one regime of post-training, not a universal explanation.