Self-Play Pretraining with Zero Data

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the bottleneck of reliance on human-generated data in pre-training by proposing a pioneering zero-data pretraining paradigm. Methodologically, synthetic data is generated via dual-model self-play, while reinforcement learning, cross-entropy loss, and universal Turing machine mechanisms are integrated to search for computable structures. An adaptive curriculum further enables the co-optimization of the generator and learner. This paradigm reframes data generation as an exploration of the entire computable space, thereby transcending the constraints of human knowledge. Experiments demonstrate that models trained under this paradigm exhibit predictable computational scaling laws on natural datasets, alongside emergent capabilities such as in-context learning and the autonomous discovery of mathematical sequences.
📝 Abstract
Advances in language modeling have been driven by scaling pretraining on ever more data. Yet, the training data is still largely curated on the model's behalf. A more general approach to pretraining would let the model learn to generate the data most useful for its own improvement. This would provide an effectively unbounded source of training data, limited by compute rather than human knowledge. We introduce Self-Play Pretraining with Zero Data, an initial proof-of-concept towards realizing this vision. Our procedure casts synthetic data generation as a search over the space of all computable structure, taking inspiration from Solomonoff induction. Starting from random initialization, two models learn in tandem: a generator proposes programs interpreted by a universal Turing machine, generating byte sequences, while a learner autoregressively predicts these byte sequences. The learner is trained with standard cross-entropy, while the generator is trained with reinforcement learning to produce sequences at the frontier of the learner's capabilities, yielding an adaptive curriculum. A universal Turing machine gives us a search space over all computable data-generating processes, imposing little domain-specific structure, and self-play searches over this space for useful training data. We test whether zero-shot performance on natural data improves predictably with self-play compute; this is a clean test of transfer since neither generator nor learner is trained on natural data. Across several natural datasets, zero-shot loss exhibits predictable scaling in compute. The models also exhibit in-context learning, and discover recognizable mathematical sequences during training.
Problem

Research questions and friction points this paper is trying to address.

Self-Play Pretraining
Zero Data
Synthetic Data Generation
Scaling Laws
Transfer Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Play Pretraining
Zero Data
Solomonoff Induction
Universal Turing Machine
Reinforcement Learning
🔎 Similar Papers