When Updating Stops Being Learning: Rethinking LLM Self-Evolution via learnable information gain

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the performance degradation phenomenon in large language model self-evolution, where capabilities initially improve but subsequently decline, as well as the limitation of existing methods that overlook system coupling. For the first time, this work formulates self-evolution as a tightly coupled system and proposes a holistic optimization framework based on learnable information gain. The framework defines information gain as the sum of KL divergence and entropy change, employing a smaller model to approximate historical data distributions and estimate the information gain of new samples via negative log-likelihood. Building upon this, an ATRI strategy is introduced to dynamically reweight samples and adaptively terminate training. Experimental results demonstrate that the proposed approach significantly delays or entirely prevents self-evolutionary degradation across mainstream benchmarks, consistently outperforming existing baselines.
📝 Abstract
Self-evolution lets large language models (LLMs) improve iteratively using their own generated data, but often suffers from self-evolution degeneration: performance improves, plateaus, then declines. Existing methods address this issue at the component level, targeting either the Questioner or the Solver, and overlook that self-evolution is a tightly coupled system. We propose a holistic framework based on learnable information gain, which measures how much novel, parameterizable information a round provides relative to the previous round. Theoretically, this gain equals the Kullback-Leibler divergence between the two rounds' data distributions plus their entropy change. Practically, it is estimated by fitting a small language model to the previous round and scoring new data via negative log-likelihood. Based on this diagnostic, we propose ATRI (Adaptive Training Regulation via Information-gain), which reweights samples within a round and halts training across rounds when information gain remains low. Experiments on popular datasets demonstrate the superiority of our proposal.
Problem

Research questions and friction points this paper is trying to address.

Self-evolution degeneration
Large language models
Information gain
Self-improvement
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Evolution
Information Gain
Large Language Models
Adaptive Training Regulation
KL Divergence
C
Chenxu Wang
Beijing University of Posts and Telecommunications
Chaozhuo Li
Chaozhuo Li
Microsoft Research Aisa
X
Xinze Shi
Beijing University of Posts and Telecommunications
S
Songyang Liu
Beijing University of Posts and Telecommunications
K
Kyrie You Wu
Beijing University of Posts and Telecommunications
Z
Ziluowen Luo
Central South University
S
Shun Zhang
Graduate School of China Academy of Engineering Physics
C
Chenxi Li
The Chinese University of Hong Kong, Shenzhen
Litian Zhang
Litian Zhang
Beihang University