Beyond Endpoint Performance: Process-Level Evaluation of Self-Evolving Agents

📅 2026-09-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出EvoPathBench,通过追踪个体能力来评估自进化代理的过程级性能,解决了仅依靠终点性能评价不全面的问题。
📝 Abstract
Self-evolving agents convert interaction feedback into persistent artifacts, such as memories or skills, which in turn guide subsequent decisions. As these artifacts are iteratively updated throughout an experience stream, the capabilities they support may evolve. Consequently, endpoint performance alone offers an incomplete view of self-evolution. Process-level evaluation is therefore essential to identify when a target capability emerges and whether later updates strengthen, preserve, or weaken it. Motivated by this, we propose \textsc{EvoPathBench}, a benchmark that tracks individual capabilities during artifact-level self-evolution. EvoPathBench fixes the base model, tools, freezes evolving artifacts at successive checkpoints, and evaluates the target capability on held-out episodes. This benchmark evaluates agent self-evolution using public trading data and calibrated trajectories. It tests three capabilities: generalization to unseen tasks, retention after unrelated learning, and rule adaptation to new evidence. Experimental results show that gains on similar unseen tasks often weaken under distribution shift, retention losses are concentrated in a minority of evolution paths, and no method achieves reliable rule adaptation. Moreover, while self-evolution enables agents to generate candidate artifacts with substantial held-out gains, the selected updates consistently fall short of realizing this potential. Together, these findings establish capability-level process evaluation as a foundation for analyzing self-evolution, identifying candidate evaluation and selection as key targets for improvement.
Problem

Research questions and friction points this paper is trying to address.

self-evolving agents
process-level evaluation
capability evolution
Innovation

Methods, ideas, or system contributions that make the work stand out.

EvoPathBench
process-level evaluation
self-evolving agents
capability tracking
artifact evolution
H
Hongqiang Lin
Zhejiang University
C
Chao Liu
Alibaba Group
X
Xiaofan Bai
Alibaba Group
X
Xuan Jin
Alibaba Group
Y
Yuhong Li
Alibaba Group
Nenggan Zheng
Nenggan Zheng
Zhejiang University
Cyborg RobotBrain Computer InterfaceComputational Ethology
X
Xipeng Cao
Alibaba Group