Score
Designs and implements infrastructure and control software to schedule, coordinate, and execute physical robot experiments and multi-robot rollouts, including synchronization, safety controls, and mechanisms for automatic scene reset and trial sequencing. Builds and operates scalable orchestration, logging, and metric-aggregation pipelines and tooling to run many trials reproducibly across robots and to monitor and analyze execution outcomes.
This study addresses the persistent gap between theoretical control performance and its practical realization in real-world robotic systems, often caused by inadequate discretization, insufficient real-time guarantees, and weak error handling in control software. For the first time from a software engineering perspective, the authors systematically analyze 184 open-source robotic controllers through code review, empirical analysis, and test evaluation, uncovering common deficiencies in application scenarios, implementation details, and verification practices. The findings reveal that most implementations fail to properly account for critical system constraints, and their testing strategies inadequately validate the theoretical assurances they claim. This work highlights a significant disconnect between implementation quality and theoretical promises, offering concrete directions and practical guidelines for developing reliable, verifiable robotic control software.
How can the openness, reproducibility, and trustworthiness of autonomous robotic scientific experiments be ensured? This paper proposes a semantic execution tracing framework built upon the AICOR Virtual Research Building platform, which semantically aligns robotic belief states with heterogeneous sensor data. By integrating digital twin technology, a deterministic execution engine, and a cloud-native architecture, the framework enables end-to-end traceable logging and cross-platform experimental reproducibility. Its key innovation is the first implementation of a robot experiment digital twin supporting semantic memory and real-time verification—rendering experimental processes transparent, reasoning auditable, and results independently verifiable. This paradigm significantly enhances the reliability and collaborative efficiency of automated scientific research, providing foundational infrastructure for autonomous systems to meaningfully contribute to scientific discovery. (149 words)
This work addresses the control performance limitations in tree-structured robotic systems caused by hierarchical data dependencies that introduce latency from perception to decision-making. To mitigate this, the authors propose FineMote, a novel framework that introduces, for the first time, a static scheduling mechanism tailored to tree-based robot models. FineMote objectifies heterogeneous low-level control logic and determines execution order statically at compile time based on the device tree, enabling low-overhead scheduling. The approach rigorously enforces deadline and priority constraints and derives a theoretical upper bound on intra-tree decision latency. Experimental evaluation on a physical robotic platform demonstrates substantial improvements in timing behavior and runtime responsiveness, confirming the framework’s effectiveness and practicality.
This work addresses the lack of closed-loop control in traditional software development lifecycles, which often fails to simultaneously ensure security, auditability, and highly reliable automation. The authors propose a deterministic autonomous control framework that models the lifecycle as a seven-stage automated pipeline, integrating Jira-based task orchestration, structured context, resource constraints, and human-review gating mechanisms to establish a secure closed loop. Key innovations include a state-contract-based collision locking mechanism, a degradation protocol for fallback operation, and a traceable control architecture. Implemented with 12,661 lines of Python code and 6,907 lines of versioned prompt specifications—including 101 exception handlers and 12 centralized locks—the system achieved a 100% success rate (95% CI [97.6%, 100%]) across 152 initial runs, producing over 795 artifacts. All 51 issues identified through adversarial review were fully resolved, with 60% of security tickets autonomously completed.
This work addresses the frequent failures of toolchains in robotic policy training and the absence of reliable evaluation and recovery mechanisms. It proposes AgenticRobotics—a backend-agnostic agentic control plane powered by large language models that dynamically orchestrates ephemeral worker nodes to establish a recoverable and verifiable “train–evaluate–refine” transactional loop. Key innovations include an evidence-gated promotion mechanism, commit-key-based crash recovery, and a signed skill repository with a verified tool registry. The system enables zero-loss, zero-duplication execution and valid decision-making at arbitrary points under unattended operation, reducing erroneous promotions from 0.005–0.021 to 0.001 and successfully detecting all six classes of artifact tampering.
This study evaluates the reliability and adaptability of large language models in executing scientific tasks within real-world physical environments, with a focus on their ability to generate executable experimental protocols and iteratively refine them based on empirical evidence. Leveraging a robotic chemistry laboratory comprising 45 modular workstations and conducting 4,608 trials, this work extends scientific agent evaluation beyond pure reasoning to encompass physical executability and evidence-driven closed-loop adaptation, introducing a quantifiable framework for assessing deployment readiness. Results reveal that only 3.3% of generated protocols were deemed executable by expert reviewers, with the best-performing system achieving a success rate of 28.1%. Most generated workflows contained no more than 30 steps and generally lacked capabilities for workflow-level replanning or methodological reconfiguration in response to experimental outcomes.
This work addresses the challenges of dependency isolation, compatibility, reproducibility, and hardware resource sharing in multi-user collaborative and heterogeneous robotic deployments. To this end, it proposes a containerized architecture tailored for robot teams operating within edge–cloud协同 environments. The architecture uniquely integrates system-level containers (LXC/LXD), ROS 2/DDS communication middleware, and a three-tier edge infrastructure—comprising infrastructure core, platform orchestration, and compute acceleration—to enable topology-aware networking, strong isolation, and controllable resource sharing. Experimental validation in a real-world robotic laboratory demonstrates that the proposed approach significantly simplifies software integration, improves resource utilization, and supports secure prototyping alongside reproducible collaborative experimentation.
This work addresses the challenge that traditional model-based testing is ill-suited for distributed robotic systems due to their high nondeterminism, dynamic reconfiguration, and inherent complexity. To overcome this limitation, the paper proposes the Scenario Specification Language (SCSL), which enables the construction of system-level tests by composing basic scenarios. The approach integrates runtime online test generation and execution with mechanisms for dynamic component joining/leaving and interface reconnection, thereby supporting automated testing and dynamic reconfiguration. The syntax and semantics of SCSL are validated through a robotic salvage mission case study, where automatically generated tests effectively demonstrate the feasibility and advantages of the proposed method.
This study addresses the challenges of high latency, unstable concurrency, and security risks faced by large language model (LLM) agents in automating asset lifecycle management within Industry 4.0. The authors propose a Plan-then-Execute architecture that generates verifiable workflow graphs and integrates a topology-aware parallel scheduling mechanism to enable controlled inference overlap while ensuring functional correctness and security. Key technical contributions include topological-sort-based multi-agent scheduling, structured context pruning, dependency-aware concurrency control, and graceful degradation under fault injection. Evaluated on the AssetOpsBench benchmark, the system reduces median end-to-end latency by 1.6× (up to 1.8× for highly parallel tasks) and cuts inference overhead by approximately 30% through context pruning, all while maintaining stable task completion rates and output quality.