π€ AI Summary
This study addresses the limitation that existing agent frameworks rely predominantly on internal knowledge for evolution, resulting in constrained exploration and passive adaptation. To overcome this, we propose ScholarEvolve, a framework inspired by the paradigm of human experts learning from scholarly literature, which automatically incorporates cutting-edge research to guide continuous optimization. Methodologically, the framework modularizes evolutionary directions and leverages meta-coding agents, topic modeling, and strategy portfolio evaluation to identify improvement pathways, while supporting the dynamic integration of newly published works to enable proactive lifelong evolution. Experimental results demonstrate that ScholarEvolve significantly improves task completion rates on the AppWorld and Tau2-Bench benchmarks, achieving performance gains of up to 14 percentage points with Qwen models.
π Abstract
Language agents are expected to solve increasingly complex tasks, creating a growing need for continual improvement. One promising approach is to evolve the agent harness, the software that governs tool use, memory management, and task execution, while keeping the underlying language model fixed. Recent methods automate this process by using a meta coding agent to modify the harness based on execution feedback. However, relying on that agent's existing knowledge and observed failures can restrict exploration and make adaptation reactive. Inspired by how human experts learn from the research literature for new solutions, we introduce ScholarEvolve, a framework that automatically draws on state-of-the-art research to guide harness evolution. ScholarEvolve organizes the harness evolution directions into functional modules and uses topic modeling to identify distinct improvement strategies for each module. It implements these strategies and evaluates their combinations to improve task performance. Moreover, the framework is designed to incorporate new publications over time, allowing research advances to drive proactive lifelong evolution. Experiments demonstrate improvements on AppWorld and Tau2-Bench. ScholarEvolve raises Qwen3.5-27B task goal completion from 49.6% to 63.6% on AppWorld Challenge, and raises GPT-5.4-mini pass@1 from 72.7% to 81.9% on Tau2-Bench Telecom.