Beyond Compression: Diagnosing How Post-Training Changes Mathematical Reasoning

πŸ“… 2026-09-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitation of the single pass@1 metric in distinguishing whether mathematical reasoning gains in large language models stem from novel solution discovery, sampling compression, or memorization. To this end, it proposes a dual-state diagnostic framework based on large-K saturation, combined with cross-surface probing via multilingual and isomorphic prompts, to comparatively analyze how distillation and GRPO post-training pathways specifically influence reasoning mechanisms. The findings reveal an essential dichotomy: easy problems rely on compression efficiency, whereas hard problems depend on extending capability ceilings. Furthermore, the results demonstrate that sufficiently off-policy distillation can elevate performance ceilings on hard problems, surpassing DeepSeek-Math, while also indicating that although English-dominant distillation improves non-English reasoning, inherent linguistic gaps persist.
πŸ“ Abstract
Post-training is central to mathematical reasoning in modern large language models (LLMs), but endpoint pass@1 alone underidentifies what has changed. Gains may reflect newly reachable solutions, cheaper sampling of latent solutions, surface robustness, or memorisation. We compare three post-training paths under a common diagnostic readout: our sufficiently trained off-policy distillation trajectories, released Qwen3 off-policy-plus-on-policy distillation endpoints, and a released DeepSeek-Math endpoint trained with Group Relative Policy Optimisation (GRPO). Our probe uses cross-surface pass@K over verbatim prompts, paraphrases, numerical isomorphisms, and translations, plus consistency, distribution-shape, and verified supervised-fine-tuning (SFT) membership analyses. We find two regimes. On easier AMC problems, large-K ceilings are near saturation, so post-training mainly compresses sample cost. On harder AIME problems, post-training expands the large-K ceiling over the base model: sufficient off-policy distillation already raises this ceiling, Qwen3 released endpoints raise it further, and DeepSeek-Math GRPO does not dominate sufficient off-policy distillation at large K. English-dominant distillation improves non-English reasoning but preserves language-tier gaps. A controlled-overfit audit finds limited sensitivity in current SFT-membership probes. Compression is one regime of post-training, not a universal explanation.
Problem

Research questions and friction points this paper is trying to address.

post-training
mathematical reasoning
large language models
compression
diagnostic evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

post-training diagnostics
mathematical reasoning
off-policy distillation
GRPO
cross-surface pass@K
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
H
Hongyang Li
University of Luxembourg
Yiming Zhu
Yiming Zhu
Phd of AI
Social computingInternet MeasurementsData Science
X
Xiao Li
Seafill Open-Source Community
C
Caesar Wu
University of Luxembourg
S
Said Mammar
UniversitΓ© Paris-Saclay
Pascal Bouvry
Pascal Bouvry
Professor of Computer Science, University of Luxembourg
OptimisationCloud/Distributed/Parallel ComputingAd Hoc networks