🤖 AI Summary
This study addresses the scarcity of reasoning training data and the low quality of synthetic data for large language models by proposing a method to generate superhuman chain-of-thought trajectories through board game self-play. The approach employs a MuZero-style network coupled with execution environment simulation to transform search processes into structured reasoning. By distilling self-play search, it enables environment-grounded supervised learning that transfers to mathematical tasks without human annotation. Experimental results demonstrate that this cross-domain reasoning transfer significantly enhances performance: Qwen3-4B achieves an average score improvement from 24.1 to 36.6 across mathematical benchmarks, alongside substantially increased game win rates, thereby validating the effectiveness of leveraging board game self-play to bolster general reasoning capabilities in language models.
📝 Abstract
Improving reasoning abilities in Large Language Models (LLMs) requires high-quality data that exposes difficult decisions, competing alternatives, and their consequences. Data scarcity is driven by the low quality of synthetic data and the cost of human labeling. We introduce Self-Play Search Distillation (SPSD), a framework for generating superhuman synthetic data via self-play of MuZero-like networks trained on board games. SPSD uses executable environments to turn search into structured reasoning problems. At each state, the expert identifies a preferred decision, plausible alternatives, plausible opponent replies, and value estimates. By converting the self-play search records into superhuman chains-of-thought, we train LLMs with environment-grounded supervision. Although trained only on self-play search records, SPSD transfers to unseen mathematics. On Qwen3-4B-Base, it raises the mean over six mathematics benchmarks from 24.1 to 36.6 while increasing the held-out-game win rate from 15% to 45%. SPSD offers an annotation-efficient way to create high-quality synthetic data for improving LLM performance in reasoning tasks.