🤖 AI Summary
This study addresses the scalability limitations of existing scientific embodied AI benchmarks, which rely on manual task engineering to convert experimental protocols into executable verification tasks. We propose SciHorizon-eLab, a novel framework that formulates scientific task construction as a compilation problem. Through semantic grounding, agent-based task synthesis, and a simulation-driven multi-stage certification mechanism, the framework automates the translation and validation of natural language protocols into embodied tasks. Leveraging this approach, we construct BenchName, a benchmark comprising 300 certified tasks. Evaluations reveal that state-of-the-art policies achieve only a 49.7% success rate, exposing significant shortcomings in current human-AI collaborative capabilities within scientific workflows.
📝 Abstract
Embodied agents offer a promising route to automating scientific experimentation, yet their progress is constrained by the lack of reliable and systematic evaluation environments. Existing simulation-based laboratory benchmarks rely heavily on manual task engineering, making it challenging to systematically compile diverse scientific protocols into executable and verifiable embodied tasks at scale. To address this challenge, we introduce SciHorizon-eLab, an agentic protocol-to-task compiler that formulates scientific embodied task construction as a compilation problem. Given a natural-language protocol of scientific experiments, SciHorizon-eLab progressively compiles laboratory protocols into semantic-preserving embodied tasks through semantic grounding, executable task synthesis, and multi-stage simulation-based certification. The system generates semantically grounded environments, executable manipulation programs, and step-level success specifications, while enabling reproducible generation of expert demonstrations and execution traces. Using this pipeline, we further construct \BenchName, a ready-to-use benchmark comprising 300 certified tasks across diverse laboratory operations. It supports HIL task execution, reproducible expert-demonstration generation, and ordered step-level evaluation. Across representative tasks, the strongest policy attains an average success rate of only 49.7%, with further evaluations revealing pronounced weaknesses in human and embodied agent coordination. We publicly release the code, benchmark data, and evaluation toolkit at https://github.com/SciHorizon-elab/SciHorizon-elab.