SAGE: A Statistical Acceptance Gate for Self-Evolving Agents

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of traditional verification gating in LLM agent self-evolution, where aggregated scores obscure performance regressions and optimization is hindered by the optimizer's curse. To overcome these challenges, this work proposes SAGE (Statistical Acceptance Gating), a method that exposes hidden regressions through item-wise paired comparisons. By incorporating one-sided hypothesis testing, SAGE rigorously filters for statistically significant improvements while penalizing degradation, thereby enabling conservative and robust skill document editing. Experiments conducted across five benchmarks and four models demonstrate that SAGE effectively reduces regression rates in 19 out of 20 scenarios—achieving a 0% regression rate on LiveMath—while consistently delivering superior final performance in all evaluated settings.
📝 Abstract
Large Language Model (LLM)-based agents increasingly self-evolve by editing a persistent skill document that encodes their workflow, tool-use rules, and decision logic. This loop has two steps, an optimizer that proposes a candidate edit and a gate that accepts or rejects it. Prior work has concentrated on the optimizer, while the gate still follows a naive rule that keeps any edit which improves an aggregate validation score. We show that this rule fails in two ways. First, it admits permanent regressions, since an edit can raise the average while breaking items the skill already solves. Second, it is vulnerable to the Optimizer's Curse, since the best observed score on a finite and noisy validation set is upward biased. To solve the above two limitations, we propose a statistical acceptance gate for self-evolving agents (SAGE). Compared with previous work, SAGE has two contributions. First, SAGE proposes a per-item paired comparison that evaluates the current skill and the edited skill on identical validation items, which exposes regressions that an aggregate score hides and penalizes them asymmetrically. Second, SAGE also employs a one-sided paired test that commits an edit only when its wins are statistically reliable against its losses, and it abstains otherwise. SAGE is a conservative refinement of the standard gate that recovers the baseline exactly at a boundary setting. It commits only a subset of the baseline's edits, filtering out those whose gains are unreliable or purchased by breaking already-solved items. Across five benchmarks and four backbone LLMs under an equal-budget protocol, SAGE lowers the regression rate in 19 of 20 settings and matches the baseline in the remaining one, for example from 36.5% to 0% on LiveMath and from 42.8% to 0% on OfficeQA with DeepSeek-V4. SAGE also attains the highest final score in all 20 settings, raising LiveMath from 34.15 to 48.78.
Problem

Research questions and friction points this paper is trying to address.

self-evolving agents
acceptance gate
performance regression
optimizer's curse
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Evolving Agents
Statistical Acceptance Gate
Paired Comparison
Optimizer's Curse
Large Language Models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yihao Wang
Peking University
Linhan Xia
Linhan Xia
University of Oklahoma
Natural Language ProcessingDeep LearningArtificial Intelligence
R
Rui Liu
Imperial College London
Z
Zhaofeng Zhang
University of Michigan
H
Hongyu Wu
University of Oklahoma
Y
Yang Yang
Xunce Technology
J
Jinglu He
Xunce Technology
Y
Yu Guo
GienTech Technology
Kai Lei
Kai Lei
Research Professor, Peking University, Shenzhen Graduate School
Future InternetData MiningBlockchain