🤖 AI Summary
This work addresses the lack of theoretical understanding of self-improvement mechanisms in large language models (LLMs). We formally define the “generation–verification gap” and develop a modular, controllable experimental framework to systematically analyze it across pretraining, post-training, and test-time inference. Leveraging mathematical modeling, cross-model-family ablation studies, and a joint technique combining self-verification, data filtering, and knowledge distillation, we demonstrate that the gap scales monotonically with pretraining compute and exhibits a well-defined boundary governed by a scaling law. Our study establishes the first theory-driven analytical paradigm for LLM self-improvement, identifies key influencing factors—including model scale, data quality, and verification fidelity—and proposes a reproducible pathway for performance enhancement. These findings provide foundational support for developing trustworthy, self-evolving AI systems.
📝 Abstract
Self-improvement is a mechanism in Large Language Model (LLM) pre-training, post-training and test-time inference. We explore a framework where the model verifies its own outputs, filters or reweights data based on this verification, and distills the filtered data. Despite several empirical successes, a fundamental understanding is still lacking. In this work, we initiate a comprehensive, modular and controlled study on LLM self-improvement. We provide a mathematical formulation for self-improvement, which is largely governed by a quantity which we formalize as the generation-verification gap. Through experiments with various model families and tasks, we discover a scaling phenomenon of self-improvement -- a variant of the generation-verification gap scales monotonically with the model pre-training flops. We also examine when self-improvement is possible, an iterative self-improvement procedure, and ways to improve its performance. Our findings not only advance understanding of LLM self-improvement with practical implications, but also open numerous avenues for future research into its capabilities and boundaries.