Self-Play Search Distillation for Large Language Model Reasoning
This study addresses the scarcity of reasoning training data and the low quality of synthetic data for large language models by proposing a method to generate superhuman chain-of-thought trajectories through board game self-play. The approach employs a MuZero-style network coupled with execution environment simulation to transform search processes into structured reasoning. By distilling self-play search, it enables environment-grounded supervised learning that transfers to mathematical tasks without human annotation. Experimental results demonstrate that this cross-domain reasoning transfer significantly enhances performance: Qwen3-4B achieves an average score improvement from 24.1 to 36.6 across mathematical benchmarks, alongside substantially increased game win rates, thereby validating the effectiveness of leveraging board game self-play to bolster general reasoning capabilities in language models.