Which Self-Improvements Should We Trust? Reliable Self-Improvement When Agents Reuse Their Benchmarks

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the problem of adaptive overfitting in recursive self-improvement, where the repeated use of fixed benchmarks yields unreliable evaluations. To mitigate this, we propose REUSE, a framework that introduces the first certified evaluation mechanism capable of handling the adaptive dependence between candidates and evaluation sets by integrating strict feedback constraints, sequential risk control, and multiple-comparison correction techniques. Theoretically, REUSE provides simultaneous error control alongside guaranteed lower bounds on cumulative improvement. Empirically, the framework reduces the spurious improvement rate from 20.7% to 0%, achieving final genuine performance comparable to the strongest baselines. These results demonstrate that REUSE enables safe and reliable self-improvement without compromising evaluation integrity.
📝 Abstract
As recursive self-improvement (RSI) rapidly advances, reliable evaluation becomes critical for guiding adaptive search. RSI typically relies on finite evaluation resources, such as fixed benchmarks, to determine which modifications are retained and what is proposed next. However, when these finite resources are repeatedly reused, new candidates are proposed based on feedback from the same evaluation set, so the search trajectory can adaptively overfit and empirical improvement may not reflect genuine population improvement on the underlying task distribution. Some existing methods account for multiple comparisons but assume that candidates are chosen independently of the evaluation set, and therefore do not control this adaptive dependence. To address this, we propose REUSE (Risk-controlled Evaluation Under Sequential Evolution), a certified evaluation and promotion framework that allows a fixed evaluation set to support repeated adaptive decisions while providing statistical guarantees. For a user-specified error level $\alpha$, with probability at least $1-\alpha$, every promoted modification is a genuine population improvement on the underlying task distribution. REUSE achieves this by strictly limiting the evaluation feedback returned to the search process and accounting for possible promotion histories within the error budget. We develop detailed statistical theory for RSI evaluation in this setting, including simultaneous error control, valid lower bounds on cumulative improvement, and a characterization of the fundamental limits of adaptive evaluation reuse. In live self-improvement experiments, REUSE commits substantially fewer false promotions than evaluation frameworks from current RSI systems and error-controlled baselines, reducing the proportion of false promotions from up to 20.7% to 0%, while achieving final true population performance comparable to the best baselines.
Problem

Research questions and friction points this paper is trying to address.

recursive self-improvement
adaptive overfitting
benchmark reuse
evaluation reliability
false promotion
Innovation

Methods, ideas, or system contributions that make the work stand out.

Recursive Self-Improvement
Adaptive Overfitting
Statistical Guarantees
Risk-controlled Evaluation
Benchmark Reuse
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xiaojing Sun
Purdue University
Y
Yuhan Zeng
Purdue University
Z
Zihua She
Purdue University
Xiao Wang
Xiao Wang
Professor of Statistics, Purdue University
Data ScienceAINonparametric StatisticsFunctional Data Analysis