Score
Designs, builds, and operates controlled experimental deployments on testbeds by installing and configuring hardware and software, integrating measurement and logging infrastructure, and setting up procedures for running experiments. Runs and manages reproducible, safe trials with baselines, collects and analyzes performance metrics, and documents the setup to ensure repeatability and reliable comparisons.
In model-based systems engineering, low experimental data reuse efficiency and excessive redundant experiments hinder digital engineering agility. To address this, this paper proposes a case-based reasoning (CBR)-driven experimental management framework that explicitly integrates domain knowledge. The framework features structured experimental metadata modeling, digital twin–enabled scenario semantic alignment, and an interpretable similarity assessment mechanism to intelligently determine whether historical experiments can be transferred to address new verification queries. Its key innovation lies in embedding domain knowledge explicitly into both the CBR retrieval and adaptation stages, thereby enabling trustworthy cross-operating-condition and cross-configuration experimental data reuse. Evaluated on an industrial-scale vehicle energy system design case, the framework reduces redundant experiments by 37% and shortens early verification cycles by 42% on average, significantly enhancing iterative efficiency in digital engineering and advancing intelligent experimental management.
This study addresses the challenge of limited reproducibility and transparency in software engineering controlled experiments, often stemming from inadequate documentation. While generic preregistration templates—such as those provided by the Open Science Framework (OSF)—exist, they fail to comprehensively address the specific needs of software engineering research. This work presents the first systematic evaluation of the OSF preregistration template’s applicability to software engineering experiments, combining literature analysis, template comparison, and cross-referencing against established software engineering experimental reporting guidelines. The findings reveal that although existing OSF templates partially satisfy methodological requirements, none fully encompass all critical elements, and their customization capabilities are constrained. Based on these insights, the paper advocates for and provides a foundation toward developing a domain-specific, standardized preregistration template tailored to software engineering, thereby filling a critical gap in the field.
This study addresses the persistent gap between theoretical control performance and its practical realization in real-world robotic systems, often caused by inadequate discretization, insufficient real-time guarantees, and weak error handling in control software. For the first time from a software engineering perspective, the authors systematically analyze 184 open-source robotic controllers through code review, empirical analysis, and test evaluation, uncovering common deficiencies in application scenarios, implementation details, and verification practices. The findings reveal that most implementations fail to properly account for critical system constraints, and their testing strategies inadequately validate the theoretical assurances they claim. This work highlights a significant disconnect between implementation quality and theoretical promises, offering concrete directions and practical guidelines for developing reliable, verifiable robotic control software.
This study addresses operational inefficiencies, poor auditability, and upgrade challenges in legacy control systems of large-scale scientific facilities—such as CERN, Diamond Light Source, and Fermilab’s ACORN project. We propose a GitOps-based modernization framework that adopts Git as the single source of truth for declarative configurations and tightly integrates containerization, Infrastructure-as-Code (IaC), and cloud-native principles to establish an automated, traceable, and version-controlled control infrastructure. Notably, this work represents the first systematic integration of modern data pipelines and AI/ML capabilities into accelerator science control systems, enabling automated configuration deployment, closed-loop runtime telemetry, and intelligent anomaly detection. Empirical evaluation demonstrates significant improvements in system reliability, maintainability, and regulatory audit compliance. The approach provides a reusable technical paradigm and engineering framework for the digital transformation of big-science facilities.
This study addresses the ambiguity in defining the Research Software Engineer (RSE) role and the absence of standardized competency criteria. Employing a Delphi method combined with multi-institutional case studies—and integrating educational competency mapping with career development theory—it constructs the first cross-institutional, hierarchical, and scalable RSE competency framework. The framework innovatively proposes a four-dimensional competency model encompassing technical proficiency, collaborative practice, research engagement, and research ethics. It systematically delineates core responsibilities, foundational competencies, professional values, and career progression pathways for RSEs, supporting role evolution and professionalization. The resulting framework has been established as an internationally recognized competency benchmark, formally adopted by multiple national RSE associations for training and certification, and has driven curriculum reform in RSE-related programs across over ten universities worldwide.
Existing agent evaluation benchmarks predominantly focus on virtual software interactions and fail to assess the multimodal interface coordination and feedback-driven parameter tuning required for scientific instrument control. This work introduces the first benchmark specifically designed for this domain, presenting a web-based, extensible, secure, and reproducible simulator suite encompassing eight instrument types and 96 subtasks that fully span the workflow from sample loading to result inspection. The benchmark supports flexible task configuration and execution-based evaluation, integrating vision-language models with a dedicated agent framework. Experimental results demonstrate that while current agents can handle structured GUI subtasks, they struggle significantly with feedback-driven operations and long-horizon workflows, thereby validating the benchmark’s necessity and its capacity to expose critical gaps in agent capabilities.
This work addresses the challenge that existing AI experimentation platforms struggle to simultaneously support rapid prototyping and governance requirements such as access control, tenant isolation, and process transparency. The authors propose and implement a governance-aware, multi-tenant AI sandbox platform featuring a layered architecture that decouples the user interface, control plane, and execution layer. The platform integrates approval workflows, audit logging, and configuration persistence mechanisms, and innovatively combines structured experimentation with cross-project reusable evaluation evidence generation, thereby establishing persistent linkages between governance decisions and experimental data. Deployed in an industry–academia collaboration setting, the platform demonstrates its effectiveness in enabling controlled collaboration, traceable experiments, and cross-project result comparison, offering a reusable reference architecture and practical insights for integrating governance into AI development environments.