Score
Designs and builds automated platforms, scripts, and control systems that execute, orchestrate, and monitor experiments or tests at scale — from robotic laboratory workflows and high‑throughput screening to browser, API, and UI test harnesses. Implements measurement and data‑collection automation, test orchestration and quality‑control pipelines, and integrates feedback/closed‑loop model control while analyzing reliability, reproducibility, and throughput of the automation.
This study addresses the challenges of laboratory automation—particularly the complexity of instrument coordination and cumbersome software configuration—that significantly hinder research efficiency. The authors propose an AI agent framework integrating large language models (LLMs) with the Experiment Orchestration System (EOS), pioneering the embedding of LLMs into the full lifecycle management of experiments. This system enables scientists to interactively create, execute, monitor, and optimize experimental protocols through natural language, complemented by a node-based visual graph editor for intuitive and synchronized protocol construction. Leveraging an agent-loop architecture alongside automated validation and error-correction mechanisms, the framework achieves a 97% success rate in translating natural language instructions into executable protocols and supports seamless switching between AI-driven and manual editing. Evaluations in simulated chemical, biological, and materials science environments demonstrate a tenfold reduction in user operations.
This work addresses the lack of closed-loop control in traditional software development lifecycles, which often fails to simultaneously ensure security, auditability, and highly reliable automation. The authors propose a deterministic autonomous control framework that models the lifecycle as a seven-stage automated pipeline, integrating Jira-based task orchestration, structured context, resource constraints, and human-review gating mechanisms to establish a secure closed loop. Key innovations include a state-contract-based collision locking mechanism, a degradation protocol for fallback operation, and a traceable control architecture. Implemented with 12,661 lines of Python code and 6,907 lines of versioned prompt specifications—including 101 exception handlers and 12 centralized locks—the system achieved a 100% success rate (95% CI [97.6%, 100%]) across 152 initial runs, producing over 795 artifacts. All 51 issues identified through adversarial review were fully resolved, with 60% of security tickets autonomously completed.
This study addresses the persistent gap between theoretical control performance and its practical realization in real-world robotic systems, often caused by inadequate discretization, insufficient real-time guarantees, and weak error handling in control software. For the first time from a software engineering perspective, the authors systematically analyze 184 open-source robotic controllers through code review, empirical analysis, and test evaluation, uncovering common deficiencies in application scenarios, implementation details, and verification practices. The findings reveal that most implementations fail to properly account for critical system constraints, and their testing strategies inadequately validate the theoretical assurances they claim. This work highlights a significant disconnect between implementation quality and theoretical promises, offering concrete directions and practical guidelines for developing reliable, verifiable robotic control software.
This work addresses the challenges researchers face in coordinating hardware control, data analysis, and experimental planning when building autonomous experimentation systems. To overcome these barriers, the authors propose a general-purpose, service-oriented autonomous experimentation platform featuring a language-agnostic, modular architecture. The platform enables users to define custom modules for hardware control, data processing, and experimental planning, with efficient inter-module communication facilitated through protobuf and gRPC. Integrated components—including a unified central control interface, automated UI generation, data management infrastructure, and experimental design tools—support closed-loop autonomous experimentation. By significantly lowering deployment complexity while enhancing flexibility and scalability, this approach allows researchers to concentrate on domain-specific scientific innovation rather than system integration overhead.
This work proposes a robotic platform integrating multi-agent collaboration and chemical perception to overcome the limitations of conventional automated chemistry systems, which rely on rigid, predefined workflows and struggle with the long-tailed distribution of experimental tasks and unconventional conditions. By enabling task decomposition, dynamic scheduling, and feedback-driven control, the platform achieves flexible adaptation across diverse experimental scenarios. Notably, it represents the first integration of a multi-agent system with real-time chemical sensing, thereby transcending the generalization barriers inherent in scripted automation. In acid–base titration experiments, the system demonstrated autonomous progress tracking, adaptive reagent dispensing, and robust end-to-end execution, validating its generalization capability and practical utility in complex, dynamic laboratory environments.
This work addresses the frequent failures of toolchains in robotic policy training and the absence of reliable evaluation and recovery mechanisms. It proposes AgenticRobotics—a backend-agnostic agentic control plane powered by large language models that dynamically orchestrates ephemeral worker nodes to establish a recoverable and verifiable “train–evaluate–refine” transactional loop. Key innovations include an evidence-gated promotion mechanism, commit-key-based crash recovery, and a signed skill repository with a verified tool registry. The system enables zero-loss, zero-duplication execution and valid decision-making at arbitrary points under unattended operation, reducing erroneous promotions from 0.005–0.021 to 0.001 and successfully detecting all six classes of artifact tampering.
This study evaluates the reliability and adaptability of large language models in executing scientific tasks within real-world physical environments, with a focus on their ability to generate executable experimental protocols and iteratively refine them based on empirical evidence. Leveraging a robotic chemistry laboratory comprising 45 modular workstations and conducting 4,608 trials, this work extends scientific agent evaluation beyond pure reasoning to encompass physical executability and evidence-driven closed-loop adaptation, introducing a quantifiable framework for assessing deployment readiness. Results reveal that only 3.3% of generated protocols were deemed executable by expert reviewers, with the best-performing system achieving a success rate of 28.1%. Most generated workflows contained no more than 30 steps and generally lacked capabilities for workflow-level replanning or methodological reconfiguration in response to experimental outcomes.
This study addresses the challenges of control design in complex industrial processes characterized by multivariable coupled dynamics by proposing an automated control strategy generation framework that integrates large language models (LLMs) with Bayesian optimization. The approach decomposes control design into structured code generation steps, ensuring physical consistency through execution-based validation and feedback-driven repair. It pioneers the automatic synthesis of decentralized PI controller architectures and their tuning environments directly from dynamic process models. Evaluated on a nonlinear gas preheater benchmark, the generated control schemes—subsequently refined via Bayesian optimization—achieve a 26.5% improvement in closed-loop performance and significantly enhance the transient response of pressure loops, thereby demonstrating the method’s effectiveness and novelty.
This work addresses the challenge of aligning natural language experimental protocols, scientific intent, and executable device instructions in automated biological experimentation by introducing ProtoPilot, a self-evolving multi-agent system. ProtoPilot leverages a hierarchical verifiable architecture, collaborative multi-agent coordination, and runtime skill library updates to automatically translate natural language protocols into executable code, establishing the first closed-loop autonomous wet-lab workflow with experimental validation. Empirical evaluation demonstrates that ProtoPilot achieves a 90.2% top-3 agreement with expert preferences, an 89.5% end-to-end protocol-to-code success rate, and an 88.24% execution success rate on the Opentrons platform—significantly outperforming baseline methods. The system successfully produced DNA constructs validated by Sanger sequencing and generated interpretable experimental outcomes.