ScienceClaw: Benchmarking Continual Self-Evolution of AI-for-Science Agents Across the Natural and Social Sciences

๐Ÿ“… 2026-10-06
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the challenge that verification outcomes of AI scientific agents rarely translate into persistent programmatic improvements, alongside the absence of continuous cross-disciplinary evaluation. We propose a fixed-parameter program self-evolution framework that unifies task solving, scientific verification, and program updating. Through multi-turn interactive workflow repair, failed trajectories are automatically transformed into reusable skills. Furthermore, we introduce the first continuous self-evolution benchmark spanning natural and social sciences, employing LLM-based agent techniques and replay verification strategies for independently reset evaluations. A comprehensive assessment system encompassing 23 disciplines is constructed to quantify scientific correctness, evolutionary gains, and transferability, significantly enhancing agentsโ€™ continual learning efficacy.
๐Ÿ“ Abstract
Large language model agents are accelerating scientific automation, yet verified executions rarely become persistent program-level improvements, and existing evaluations do not examine this process across sequential tasks in both the natural and social sciences. We formalize ScienceClaw as fixed-parameter program self-evolution that unifies task solving, scientific verification, and program updates. ScienceClaw-Eval spans 23 disciplines and measures scientific correctness, evolutionary gain, retention, cross-dataset transfer, and evolution cost through sequential streams and independent reset evaluation. Our framework repairs executable workflows through multi-turn interaction, converts re-execution-verified failure--success trajectories into linked Skill and Operator candidates, and retains an update only when source-task replay reproduces the repair and independent scientific tasks improve. Code is available at https://github.com/beita6969/ScienceClaw.
Problem

Research questions and friction points this paper is trying to address.

Continual Self-Evolution
AI-for-Science Agents
Program-level Improvement
Benchmarking
Innovation

Methods, ideas, or system contributions that make the work stand out.

Continual Self-Evolution
Program-level Improvement
AI-for-Science Agents
Skill and Operator Extraction
Multi-disciplinary Benchmarking