research testbed orchestration

Designs and implements orchestration systems and workflows that provision, configure, and manage distributed testbed resources (compute, network, and storage), including automated deployment of software stacks and runtime environments. Builds experiment execution and scheduling pipelines to run distributed benchmarks and workloads, collect outputs and logs, and ensure reproducibility, fault handling, and lifecycle management across heterogeneous remote resources.

researchtestbedorchestration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$255K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.

Autonomous AgentsHigh Performance ComputingJob Specification Translation

This work addresses the limitations of traditional workflow platforms, which rely on static, pre-defined processes and struggle to accommodate the dynamic data integration demands of distributed systems. To overcome this, the authors propose a configuration-driven runtime orchestration framework that dynamically constructs execution graphs at request time through dependency-aware scheduling and parallel task execution, thereby circumventing the constraints of fixed workflows. This approach enables rapid adaptation to evolving integration scenarios without requiring system redeployment, significantly reducing latency. Empirical evaluation in a real-world Customer 360 enterprise use case demonstrates that the framework offers substantial advantages in flexibility, scalability, and efficient data aggregation compared to conventional solutions.

distributed systemsdynamic data retrievalheterogeneous integrations

This work addresses the persistent challenge of inconsistent development and execution environments faced by researchers operating across heterogeneous computing platforms—ranging from laptops and workstations to supercomputers and cloud infrastructures. To overcome this, the authors propose a modular and portable software ecosystem featuring a unified command-line interface that enables seamless orchestration and execution of scientific workflows. The system ensures cross-platform consistency, reproducibility, and scalability, thereby streamlining computational research across diverse hardware configurations. Its practical efficacy has been demonstrated through successful integration into the plan4res project under the European Union’s Horizon 2020 initiative, where it effectively supported complex, large-scale scientific workflows in varied computing environments.

computational workflowsportablereproducible

iDDS: Intelligent Distributed Dispatch and Scheduling for Workflow Orchestration

Oct 03, 2025
WG
Wen Guan
🏛️ Brookhaven National Laboratory | University of Texas at Arlington | University of Pittsburgh

To address the challenge of efficiently orchestrating and intelligently managing complex, dynamic workflows in large-scale distributed scientific computing, this paper proposes an integrated intelligent workflow system that unifies task scheduling, data movement, and adaptive decision-making. The system supports data-aware execution, conditional logic, and programmable directed acyclic graphs (DAGs), operating in both template-driven and “function-as-a-task” modes. It adopts a modular, message-driven architecture and deeply integrates mainstream middleware—including PanDA and Rucio—while incorporating distributed hyperparameter optimization and AI-assisted modeling. Its cross-experiment, cross-platform design significantly enhances scalability and reproducibility. Deployed in major scientific projects—including ATLAS, the Rubin Observatory, and the Electron-Ion Collider—the system enables high-throughput execution of heterogeneous tasks and reduces operational overhead by over 30%.

Integrating data-aware execution with conditional logic automationOrchestrating large-scale distributed scientific computing workflowsUnifying workload scheduling and data movement across infrastructures

This work addresses the growing challenge in modern scientific research where complex infrastructure management, authentication, and deployment processes divert focus from core scientific discovery. To this end, the authors propose Sci-Orchestra, a cloud-native, Kubernetes-based hierarchical orchestration framework that automates experimental workflows through API-driven mechanisms, enabling secure authentication, resource scheduling, and scalable deployment across heterogeneous high-performance computing environments. Sci-Orchestra introduces an autonomous service marketplace to facilitate cross-institutional collaboration and adopts a “black-box” interoperability model that promotes integration of academic, industrial, and research tools while safeguarding intellectual property. By significantly lowering technical barriers, the framework accelerates the transition of research prototypes into production-grade applications, thereby advancing the emerging paradigm of Science as a Service (SciaaS).

cross-institutional collaborationHPC interoperabilityinfrastructure management

Latest Papers

What's happening recently
View more

This work proposes the first large language model (LLM)-driven agent framework for autonomous, end-to-end management of high-performance computing (HPC) applications in cloud environments. Addressing the heavy reliance on manual intervention and the lack of intelligent decision-making in traditional HPC cloud deployment, the framework enables automated multi-platform container construction, Kubernetes-based orchestration, cross-instance performance optimization, and adaptive elastic scaling policy generation. By integrating LLM-powered agents into HPC cloud workflow orchestration, this study establishes a novel paradigm of automation and self-adaptation. Experimental evaluation across four representative HPC applications demonstrates that the system achieves expert-level linear scalability, substantially reduces job completion time, and yields actionable best practices for collaborative agent design in HPC contexts.

Agentic OrchestrationCloud ComputingHPC Applications

This work addresses the challenges of workflow task composition in high-throughput, petabyte-scale data processing environments, where resource heterogeneity and execution overhead significantly impact performance. The authors propose a hybrid task composition strategy that dynamically balances task independence against execution grouping, formulated within a multi-objective optimization framework to achieve Pareto-optimal trade-offs among throughput, I/O cost, and CPU efficiency. Leveraging workflow DAG modeling and high-dimensional parameter space simulation, the approach enables policy-driven automated synthesis of workflows. Experimental results demonstrate that the proposed strategy achieves up to a 3.8× improvement in throughput and reduces network overhead by as much as 14.9× compared to baseline methods, offering a scalable workflow synthesis framework for extreme-scale scientific computing.

extreme-scale data processingHigh-Throughput Computingresource utilization

This work addresses the orchestration bottlenecks faced by ultra-large-scale Sim-AI workflows on leadership-class supercomputers, which arise from task heterogeneity and extreme ensemble sizes. To overcome these challenges, the authors propose EnsembleLauncher, a recursively hierarchical and fully decentralized workflow orchestrator that introduces a decentralized control plane and a programmable scheduling policy interface, thereby surpassing conventional tools in both scalability and scheduling flexibility. Experiments on the Aurora supercomputer demonstrate that EnsembleLauncher can efficiently schedule system-wide resources to support up to 8 million serial tasks, achieving more than a fourfold performance improvement over state-of-the-art alternatives. Furthermore, it significantly enhances resource utilization for workloads with high task variance and active learning pipelines.

exascaleorchestration bottlenecksscalability

This work addresses the challenge of automated resource discovery and scheduling across heterogeneous, multi-institutional computing environments—including cloud, edge, and high-performance computing (HPC) infrastructures—in agent-based systems. The authors propose a hierarchical dynamic agent architecture in which secretary agents concurrently and asynchronously perform resource probing, negotiation, and task dispatching. By integrating an asynchronous negotiation protocol, agent-driven resource categorization, and a dynamic scheduling algorithm, the framework enables highly scalable, cross-infrastructure automation. Evaluated on a testbed comprising 51 real and simulated resource providers, the approach achieved a negotiation accuracy of 87.71% over 19,973 negotiation rounds and 6,952 task selections, with task selection costs comparable to conventional strategies, thereby significantly enhancing scheduling efficiency and adaptability.

agentic scienceheterogeneous systemshierarchical architecture

Hot Scholars

RG

Raghav Gupta

EvolutionaryScale
Natural Language Processing
SK

Sagar Karandikar

Assistant Professor in EECS, UC Berkeley
Computer ArchitectureComputer SystemsVLSI