experiment orchestration and management

Designs and implements systems and pipelines that coordinate, schedule, execute, monitor, and record experiments or trials across software and hardware components, including resource provisioning, distributed execution, logging, and result collection. Builds automation for the experiment lifecycle—versioning and provenance tracking, reproducibility, hyperparameter and batch search, monitoring and failure recovery, metadata management, and tooling/APIs for running, querying, and analyzing experimental runs.

experimentorchestrationandmanagement

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.21
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$225K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing autonomous research systems struggle to effectively accumulate and leverage trial-and-error experience, often leading to overstatement of weak evidence, repeated failures, and memory loss. This work proposes the Sibyl-AutoResearch framework, which incorporates a scientific trial-and-error harness mechanism enabling agents to conduct bounded experiments, explicitly store both positive and negative outcomes, and feed this experiential knowledge back into subsequent research phases. The framework introduces two auditable transformation units—trial-to-behavior and trial-to-harness-behavior—and integrates a file-persistence architecture that explicitly models states, roles, memories, gating mechanisms, and artifact trajectories, thereby supporting fully traceable self-evolution and self-repair. Experimental retrospection identified eight high-confidence transformation events (median delay of one round), successfully intercepting or rectifying five common failure modes, including redundant results and outdated data.

autonomous researchresearch judgmentscientific workflow

This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.

Autonomous AgentsHigh Performance ComputingJob Specification Translation

This work addresses the challenges researchers face in coordinating hardware control, data analysis, and experimental planning when building autonomous experimentation systems. To overcome these barriers, the authors propose a general-purpose, service-oriented autonomous experimentation platform featuring a language-agnostic, modular architecture. The platform enables users to define custom modules for hardware control, data processing, and experimental planning, with efficient inter-module communication facilitated through protobuf and gRPC. Integrated components—including a unified central control interface, automated UI generation, data management infrastructure, and experimental design tools—support closed-loop autonomous experimentation. By significantly lowering deployment complexity while enhancing flexibility and scalability, this approach allows researchers to concentrate on domain-specific scientific innovation rather than system integration overhead.

autonomous experimentationexperimental workflowlaboratory automation

Reasonable Experiments in Model-Based Systems Engineering

Sep 12, 2025
JC
Johan Cederbladh
🏛️ Mälardalen University | Eindhoven University of Technology | Stellenbosch University | IT University of Copenhagen | University of Oslo | Universidade Federal Rural de Pernambuco | University of Antwerp

In model-based systems engineering, low experimental data reuse efficiency and excessive redundant experiments hinder digital engineering agility. To address this, this paper proposes a case-based reasoning (CBR)-driven experimental management framework that explicitly integrates domain knowledge. The framework features structured experimental metadata modeling, digital twin–enabled scenario semantic alignment, and an interpretable similarity assessment mechanism to intelligently determine whether historical experiments can be transferred to address new verification queries. Its key innovation lies in embedding domain knowledge explicitly into both the CBR retrieval and adaptation stages, thereby enabling trustworthy cross-operating-condition and cross-configuration experimental data reuse. Evaluated on an industrial-scale vehicle energy system design case, the framework reduces redundant experiments by 37% and shortens early verification cycles by 42% on average, significantly enhancing iterative efficiency in digital engineering and advancing intelligent experimental management.

Deciding if existing experiments can answer new engineering questionsIntelligently reusing experiment-related data to avoid redundant experimentsManaging experimental configuration metadata and results efficiently

This work addresses the lack of closed-loop control in traditional software development lifecycles, which often fails to simultaneously ensure security, auditability, and highly reliable automation. The authors propose a deterministic autonomous control framework that models the lifecycle as a seven-stage automated pipeline, integrating Jira-based task orchestration, structured context, resource constraints, and human-review gating mechanisms to establish a secure closed loop. Key innovations include a state-contract-based collision locking mechanism, a degradation protocol for fallback operation, and a traceable control architecture. Implemented with 12,661 lines of Python code and 6,907 lines of versioned prompt specifications—including 101 exception handlers and 12 centralized locks—the system achieved a 100% success rate (95% CI [97.6%, 100%]) across 152 initial runs, producing over 795 artifacts. All 51 issues identified through adversarial review were fully resolved, with 60% of security tickets autonomously completed.

Autonomous Software DevelopmentBacklog OrchestrationClosed-Loop Control

Latest Papers

What's happening recently
View more

This work addresses the heavy reliance on expert knowledge in designing and debugging scientific workflows, a challenge exacerbated by existing large language model approaches that directly generate code without ensuring transparency, reproducibility, or seamless system integration. To overcome these limitations, we propose an AI-assisted scientific workflow management framework that decouples user intent from implementation through a structured specification phase, enabling specification-driven workflow generation and validation. We further introduce a multi-layer debugging agent powered by large language models to automate error diagnosis and correction. By deeply integrating with the Pegasus workflow system via the Model Context Protocol (MCP), our approach supports end-to-end workflow lifecycle management. Empirical evaluation demonstrates successful generation and execution of federated learning medical imaging workflows comprising thousands of tasks, substantially reducing debugging effort and empowering non-expert users to construct complex workflows adhering to expert-level design patterns.

debugginglarge language modelsreproducibility

This work addresses the persistent challenge of inconsistent development and execution environments faced by researchers operating across heterogeneous computing platforms—ranging from laptops and workstations to supercomputers and cloud infrastructures. To overcome this, the authors propose a modular and portable software ecosystem featuring a unified command-line interface that enables seamless orchestration and execution of scientific workflows. The system ensures cross-platform consistency, reproducibility, and scalability, thereby streamlining computational research across diverse hardware configurations. Its practical efficacy has been demonstrated through successful integration into the plan4res project under the European Union’s Horizon 2020 initiative, where it effectively supported complex, large-scale scientific workflows in varied computing environments.

computational workflowsportablereproducible

This study evaluates the reliability and adaptability of large language models in executing scientific tasks within real-world physical environments, with a focus on their ability to generate executable experimental protocols and iteratively refine them based on empirical evidence. Leveraging a robotic chemistry laboratory comprising 45 modular workstations and conducting 4,608 trials, this work extends scientific agent evaluation beyond pure reasoning to encompass physical executability and evidence-driven closed-loop adaptation, introducing a quantifiable framework for assessing deployment readiness. Results reveal that only 3.3% of generated protocols were deemed executable by expert reviewers, with the best-performing system achieving a success rate of 28.1%. Most generated workflows contained no more than 30 steps and generally lacked capabilities for workflow-level replanning or methodological reconfiguration in response to experimental outcomes.

evidence-driven replanninglong-horizon planningphysical executability

Hot Scholars

MC

Matteo Cinelli

Assistant Professor @Sapienza University of Rome
Data ScienceNetwork ScienceSocial MediaComputational Social Science
ZM

Zhipeng Ma

Southwest Jiaotong University
Data-Centric AILarge Language ModelHuman Mobility
BN

Bo Nørregaard Jørgensen

Professor, PhD., Head of Center for Energy Informatics, University of Southern Denmark
Energy InformaticsEnergy-ecosystemsAI AgentsMulti-agent systems