🤖 AI Summary
This study addresses the challenge of long-horizon manipulation for embodied intelligence in chemical experiments, where real-world data scarcity and inadequate simulation fidelity hinder adherence to strict procedural constraints. To overcome this, the authors propose a protocol-driven, high-fidelity simulation framework coupled with a self-improving task synthesis mechanism. By integrating large language model-based protocol parsing with reinforcement learning for iterative scene configuration optimization, the approach automatically generates multi-step expert trajectories compliant with experimental protocols. Furthermore, a hierarchical evaluation benchmark spanning from atomic operations to full workflows is constructed. The method successfully synthesizes reliable expert demonstrations exceeding ten steps, significantly enhancing both training efficiency and generalization capabilities of agents in complex chemical tasks, thereby advancing the development of automated intelligent laboratories.
📝 Abstract
Wet-lab experimentation serves as the gold standard for hypothesis verification in scientific discovery; yet it is inherently labor-intensive, costly, and safety-critical. Embodied agents hold the promise of automating these tedious workflows, but their development is hindered by the scarcity of real-world training data. While simulation offers a scalable alternative for producing demonstrations, current methods primarily target relatively short-horizon tasks with loosely structured interactions, failing to meet the strict procedural constraints and fine-grained manipulation demands of chemical experiments. To bridge this gap, we introduce \textbf{RoboChemGym}, a framework that autonomously generates high-fidelity manipulation demonstrations aligned with real-world experiment protocols, featuring a \textit{self-improving task synthesis} mechanism to iteratively refine task execution and scene configurations, enabling the reliable generation of expert trajectories for complex, multi-object protocols exceeding 10 interaction steps. Furthermore, we introduce a hierarchical benchmark that systematically assesses performance across varying granularities, spanning from atomic operations to full-cycle experimental workflows. RoboChemGym sets a scalable paradigm for the automated data synthesis and capability evaluation of embodied agents in intricate chemical tasks, serving as a critical stepping stone toward fully intelligent laboratories.