🤖 AI Summary
This study addresses the performance degradation phenomenon in large language model self-evolution, where capabilities initially improve but subsequently decline, as well as the limitation of existing methods that overlook system coupling. For the first time, this work formulates self-evolution as a tightly coupled system and proposes a holistic optimization framework based on learnable information gain. The framework defines information gain as the sum of KL divergence and entropy change, employing a smaller model to approximate historical data distributions and estimate the information gain of new samples via negative log-likelihood. Building upon this, an ATRI strategy is introduced to dynamically reweight samples and adaptively terminate training. Experimental results demonstrate that the proposed approach significantly delays or entirely prevents self-evolutionary degradation across mainstream benchmarks, consistently outperforming existing baselines.
📝 Abstract
Self-evolution lets large language models (LLMs) improve iteratively using their own generated data, but often suffers from self-evolution degeneration: performance improves, plateaus, then declines. Existing methods address this issue at the component level, targeting either the Questioner or the Solver, and overlook that self-evolution is a tightly coupled system. We propose a holistic framework based on learnable information gain, which measures how much novel, parameterizable information a round provides relative to the previous round. Theoretically, this gain equals the Kullback-Leibler divergence between the two rounds' data distributions plus their entropy change. Practically, it is estimated by fitting a small language model to the previous round and scoring new data via negative log-likelihood. Based on this diagnostic, we propose ATRI (Adaptive Training Regulation via Information-gain), which reweights samples within a round and halts training across rounds when information gain remains low. Experiments on popular datasets demonstrate the superiority of our proposal.