Score
Design and build automated experimentation systems that use large language models as agents to propose hypotheses and training edits, generate or modify training code, submit and monitor high-performance compute training jobs, and parse evaluation outputs. Implement the orchestration pieces—prompt/tool integration, experiment tracking, result analysis and iteration logic, and failure handling—so experiments can be launched, evaluated, and refined with minimal human intervention.
This paper addresses the paradigm shift arising from large language models (LLMs) evolving from research aids to autonomous scientific agents. Method: Through a systematic literature review and interdisciplinary analysis—integrating philosophy of science, human-AI collaboration theory, and AI governance—we propose the first three-tiered autonomy framework for LLMs in scientific discovery (Tool → Analyst → Scientist), rigorously delineating capability boundaries and evolutionary trajectories across the full scientific workflow: hypothesis generation, experimental design, and self-reflection. We anchor our conceptual architecture and technical roadmap in scientific methodology. Contribution/Results: The work delivers the field’s first comprehensive survey, establishes foundational theoretical grounding for autonomous scientific agents, and releases Awesome-LLM-Scientific-Discovery—an open-source knowledge repository supporting both theoretical advancement and practical development.
This study addresses the low automation level and high human dependency in scientific research workflows by proposing an Autonomous Simulation Agent (ASA) framework tailored for long-duration simulation tasks. Methodologically, the ASA integrates prompt engineering, automated code generation, remote high-performance computing (HPC) job scheduling, and multi-stage workflow orchestration, featuring a dynamically self-verifying architecture. A novel local-attention–global-supervision coordination mechanism enables 20 rounds of fully autonomous, human-free iteration. Evaluated on polymer chain conformational sampling, ASA-GPT-4o achieves near 100% task completion rate and sustains stable end-to-end operation across 20 consecutive cycles. The framework significantly enhances research efficiency, operational reliability, and experimental reproducibility, advancing the automation and robustness of computational science workflows.
This work addresses the low research efficiency and poor reproducibility prevalent in machine learning (ML) studies by proposing the first LLM-based autonomous scientific research framework. Methodologically, it introduces a three-stage collaborative paradigm—comprising IdeaAgent, ExperimentAgent, and ValidateAgent—integrating retrieval-augmented generation (RAG), dynamic code synthesis, experimental workflow orchestration, and secure sandboxed execution. The framework fully automates the ML research lifecycle: from hypothesis generation and adaptive selection of models/datasets to prototype code generation and closed-loop validation. Its key contribution lies in establishing the first human-in-the-loop, end-to-end ML research pipeline enabled by LLM agents, requiring no manual coding intervention. Evaluated across five representative ML tasks, the framework successfully generated and executed valid experiments, significantly improving research productivity and result reproducibility. This demonstrates the feasibility of LLM agents in driving substantive, reproducible scientific innovation in ML.
To address high manual effort, delayed response, and incomplete coverage in industrial software test maintenance, this paper proposes and empirically validates two novel multi-agent architectures leveraging large language models (LLMs) to accurately identify test cases requiring maintenance—and to assist in their automated addition, deletion, or modification—following code changes. The study systematically distills critical triggering conditions and practical deployment constraints for LLMs in industrial settings, and conducts empirical evaluation using real-world data from Ericsson AB. Results demonstrate significant reductions in manual test maintenance effort, improved timeliness of issue response, and enhanced test coverage completeness. This work represents the first application of multi-agent collaboration to industrial-scale test maintenance, establishing a reusable methodological framework and empirical foundation for deploying LLMs in high-reliability software engineering contexts.
This work addresses the current lack of open-source infrastructure capable of efficiently training and evaluating large-scale agents on complex tasks such as software engineering and computer operation. To this end, we propose a three-service decoupled architecture tailored for agent-environment interaction workloads, which separates the system into three independent services—model, agent, and environment—enabling fine-grained task scheduling, dynamic resource allocation, and unified interface communication. This design allows each component to scale independently and configure resources flexibly, significantly improving training efficiency and resource utilization. Experimental results demonstrate that the system can stably support tens of thousands of concurrent agent tasks, thereby filling a critical gap in infrastructure for large-scale agent training.
Current language agents lack rigorous evaluation and exhibit questionable task generalization in data-driven scientific discovery. Method: We introduce SciBench—the first rigorous benchmark for evaluating language agents in science—comprising 102 authentic research tasks across four disciplines, with standardized executable Python program outputs. We propose a novel three-dimensional evaluation framework assessing (i) program executability, (ii) result correctness, and (iii) computational cost. Task construction employs a paradigm combining real-paper extraction and domain-expert co-verification, augmented by a dual mechanism to prevent data contamination. Contribution/Results: Experiments on five state-of-the-art LLMs show that independent success rates across three evaluation frameworks peak at only 32.4%; integrating expert knowledge marginally improves performance to 34.3%. These results demonstrate that existing language agents remain substantially short of enabling end-to-end automated scientific discovery.
This work proposes the first large language model (LLM)-driven agent framework for autonomous, end-to-end management of high-performance computing (HPC) applications in cloud environments. Addressing the heavy reliance on manual intervention and the lack of intelligent decision-making in traditional HPC cloud deployment, the framework enables automated multi-platform container construction, Kubernetes-based orchestration, cross-instance performance optimization, and adaptive elastic scaling policy generation. By integrating LLM-powered agents into HPC cloud workflow orchestration, this study establishes a novel paradigm of automation and self-adaptation. Experimental evaluation across four representative HPC applications demonstrates that the system achieves expert-level linear scalability, substantially reduces job completion time, and yields actionable best practices for collaborative agent design in HPC contexts.
This work addresses the challenges faced by large language model (LLM) agents operating over flat tool registries—namely, combinatorial explosion in decision space, context saturation, and degraded routing accuracy. To overcome these limitations, the authors propose a skill-tree-based hierarchical architecture that separates routing logic at internal nodes from execution at leaf nodes. Inspired by pushdown automata, the framework incorporates a LIFO stack-frame memory model and a lazy capability discovery mechanism, enabling isolated execution paths and scalable context management. The approach supports manifest-driven single-step execution loops and formal state modeling, significantly improving routing accuracy while reducing memory footprint and prompt costs under conditions of tool proliferation, multi-step workflows, and prompt exposure. This design meets enterprise-grade requirements for isolation and scalability.
Existing benchmarks struggle to capture the stochasticity and adaptability of large language model agents in autonomous model discovery. This work proposes an experimental evaluation paradigm that treats agents as stochastic model discovery operators, systematically examining—through controlled experiments—how task design, objectives, data, and reasoning effort jointly influence output quality, cost, latency, and process complexity. By integrating regression modeling, statistical inference, and utility-aligned decomposition techniques, we evaluate coding agents such as Codex and Claude Code on the WordCraft platform. Our analysis reveals that reasoning effort exerts a dominant and interpretable influence on both performance and cost, thereby demonstrating the framework’s effectiveness and analytical insight.