๐ค AI Summary
This work addresses the challenge that verification outcomes of AI scientific agents rarely translate into persistent programmatic improvements, alongside the absence of continuous cross-disciplinary evaluation. We propose a fixed-parameter program self-evolution framework that unifies task solving, scientific verification, and program updating. Through multi-turn interactive workflow repair, failed trajectories are automatically transformed into reusable skills. Furthermore, we introduce the first continuous self-evolution benchmark spanning natural and social sciences, employing LLM-based agent techniques and replay verification strategies for independently reset evaluations. A comprehensive assessment system encompassing 23 disciplines is constructed to quantify scientific correctness, evolutionary gains, and transferability, significantly enhancing agentsโ continual learning efficacy.
๐ Abstract
Large language model agents are accelerating scientific automation, yet verified executions rarely become persistent program-level improvements, and existing evaluations do not examine this process across sequential tasks in both the natural and social sciences. We formalize ScienceClaw as fixed-parameter program self-evolution that unifies task solving, scientific verification, and program updates. ScienceClaw-Eval spans 23 disciplines and measures scientific correctness, evolutionary gain, retention, cross-dataset transfer, and evolution cost through sequential streams and independent reset evaluation. Our framework repairs executable workflows through multi-turn interaction, converts re-execution-verified failure--success trajectories into linked Skill and Operator candidates, and retains an update only when source-task replay reproduces the repair and independent scientific tasks improve. Code is available at https://github.com/beita6969/ScienceClaw.