Procedural Environment Generation for Tool-Use Agents

📅 2025-05-21
🏛️ arXiv.org
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing tool-augmented agents face a scarcity of high-quality training data for online reinforcement learning (RL); synthetic data typically lacks interactivity and compositional structure. Method: We propose RandomWorld, the first pipeline to programmatically generate tool-use trajectories featuring multi-step interactions and cross-tool compositionality—overcoming the static and isolated nature of conventional synthetic data. Our approach integrates programmable environment modeling, supervised fine-tuning (SFT), and Proximal Policy Optimization (PPO)-based online RL into an end-to-end training framework. Contributions/Results: On the NESTFUL benchmark, our method achieves new state-of-the-art (SOTA) performance on two core metrics. Crucially, downstream task performance scales consistently with synthetic data volume—providing the first empirical evidence that high-performance tool-using agents can be trained exclusively on synthetic data.

Technology Category

Natural Language Processing: Code Generation / Program Synthesis from Natural LanguageMultiagent Systems: Adversarial AgentsMachine Learning: Imitation Learning & Inverse Reinforcement Learning

Application Category

Search and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationEconomics, Online Markets and Human Computation: Economic ramifications for generative AI infrastructure and applicationsSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
Although the power of LLM tool-use agents has ignited a flurry of recent research in this area, the curation of tool-use training data remains an open problem$-$especially for online RL training. Existing approaches to synthetic tool-use data generation tend to be non-interactive, and/or non-compositional. We introduce RandomWorld, a pipeline for the procedural generation of interactive tools and compositional tool-use data. We show that models tuned via SFT and RL on synthetic RandomWorld data improve on a range of tool-use benchmarks, and set the new SoTA for two metrics on the NESTFUL dataset. Further experiments show that downstream performance scales with the amount of RandomWorld-generated training data, opening up the possibility of further improvement through the use of entirely synthetic data.
Problem

Research questions and friction points this paper is trying to address.

Synthetic tool-use training data curation remains an open problem
Existing approaches are non-interactive and non-compositional for tool generation
Online RL training lacks effective procedural environment generation methods
Innovation

Methods, ideas, or system contributions that make the work stand out.

Procedural generation of interactive tools and data
Synthetic data pipeline for tool-use agent training
Improves tool-use benchmarks via SFT and RL
🔎 Similar Papers
No similar papers found.
M
Michael Sullivan
Department of Language Science and Technology, Saarland Informatics Campus, Saarland University, Saarbrücken, Germany
M
Mareike Hartmann
Department of Language Science and Technology, Saarland Informatics Campus, Saarland University, Saarbrücken, Germany
Alexander Koller
Alexander Koller
Professor of Computational Linguistics, Saarland University, Saarland Informatics Campus
Computational linguisticsartificial intelligence