🤖 AI Summary
To address the limited generalization capability of LLM-based agents in specialized domains such as life sciences and medicine—stemming from their reliance on manually pre-written tools—this paper introduces ToolMaker: the first end-to-end framework for automatically constructing LLM-callable tools from scientific code repositories. Its core innovation is a closed-loop, self-correcting paradigm for fully automated tool generation, integrating multi-step reasoning–driven agent orchestration, automatic dependency installation, iterative code generation and debugging, and unit-test–driven robustness validation. Evaluated on a benchmark of 15 complex tasks spanning medical and non-medical domains, ToolMaker achieves an 80% accuracy rate, substantially outperforming existing software-engineering–oriented LLM agents. It represents the first system capable of autonomously transforming research code into production-ready, LLM-executable tools.
📝 Abstract
Tool use has turned large language models (LLMs) into powerful agents that can perform complex multi-step tasks by dynamically utilising external software components. However, these tools must be implemented in advance by human developers, hindering the applicability of LLM agents in domains which demand large numbers of highly specialised tools, like in life sciences and medicine. Motivated by the growing trend of scientific studies accompanied by public code repositories, we propose ToolMaker, a novel agentic framework that autonomously transforms papers with code into LLM-compatible tools. Given a short task description and a repository URL, ToolMaker autonomously installs required dependencies and generates code to perform the task, using a closed-loop self-correction mechanism to iteratively diagnose and rectify errors. To evaluate our approach, we introduce a benchmark comprising 15 diverse and complex computational tasks spanning both medical and non-medical domains with over 100 unit tests to objectively assess tool correctness and robustness. ToolMaker correctly implements 80% of the tasks, substantially outperforming current state-of-the-art software engineering agents. ToolMaker therefore is a step towards fully autonomous agent-based scientific workflows.