Score
Designs and implements orchestration systems and workflows that provision, configure, and manage distributed testbed resources (compute, network, and storage), including automated deployment of software stacks and runtime environments. Builds experiment execution and scheduling pipelines to run distributed benchmarks and workloads, collect outputs and logs, and ensure reproducibility, fault handling, and lifecycle management across heterogeneous remote resources.
This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.
This work addresses the limitations of traditional workflow platforms, which rely on static, pre-defined processes and struggle to accommodate the dynamic data integration demands of distributed systems. To overcome this, the authors propose a configuration-driven runtime orchestration framework that dynamically constructs execution graphs at request time through dependency-aware scheduling and parallel task execution, thereby circumventing the constraints of fixed workflows. This approach enables rapid adaptation to evolving integration scenarios without requiring system redeployment, significantly reducing latency. Empirical evaluation in a real-world Customer 360 enterprise use case demonstrates that the framework offers substantial advantages in flexibility, scalability, and efficient data aggregation compared to conventional solutions.
This work addresses the persistent challenge of inconsistent development and execution environments faced by researchers operating across heterogeneous computing platforms—ranging from laptops and workstations to supercomputers and cloud infrastructures. To overcome this, the authors propose a modular and portable software ecosystem featuring a unified command-line interface that enables seamless orchestration and execution of scientific workflows. The system ensures cross-platform consistency, reproducibility, and scalability, thereby streamlining computational research across diverse hardware configurations. Its practical efficacy has been demonstrated through successful integration into the plan4res project under the European Union’s Horizon 2020 initiative, where it effectively supported complex, large-scale scientific workflows in varied computing environments.
To address the challenge of efficiently orchestrating and intelligently managing complex, dynamic workflows in large-scale distributed scientific computing, this paper proposes an integrated intelligent workflow system that unifies task scheduling, data movement, and adaptive decision-making. The system supports data-aware execution, conditional logic, and programmable directed acyclic graphs (DAGs), operating in both template-driven and “function-as-a-task” modes. It adopts a modular, message-driven architecture and deeply integrates mainstream middleware—including PanDA and Rucio—while incorporating distributed hyperparameter optimization and AI-assisted modeling. Its cross-experiment, cross-platform design significantly enhances scalability and reproducibility. Deployed in major scientific projects—including ATLAS, the Rubin Observatory, and the Electron-Ion Collider—the system enables high-throughput execution of heterogeneous tasks and reduces operational overhead by over 30%.
This work addresses the growing challenge in modern scientific research where complex infrastructure management, authentication, and deployment processes divert focus from core scientific discovery. To this end, the authors propose Sci-Orchestra, a cloud-native, Kubernetes-based hierarchical orchestration framework that automates experimental workflows through API-driven mechanisms, enabling secure authentication, resource scheduling, and scalable deployment across heterogeneous high-performance computing environments. Sci-Orchestra introduces an autonomous service marketplace to facilitate cross-institutional collaboration and adopts a “black-box” interoperability model that promotes integration of academic, industrial, and research tools while safeguarding intellectual property. By significantly lowering technical barriers, the framework accelerates the transition of research prototypes into production-grade applications, thereby advancing the emerging paradigm of Science as a Service (SciaaS).
This work proposes the first large language model (LLM)-driven agent framework for autonomous, end-to-end management of high-performance computing (HPC) applications in cloud environments. Addressing the heavy reliance on manual intervention and the lack of intelligent decision-making in traditional HPC cloud deployment, the framework enables automated multi-platform container construction, Kubernetes-based orchestration, cross-instance performance optimization, and adaptive elastic scaling policy generation. By integrating LLM-powered agents into HPC cloud workflow orchestration, this study establishes a novel paradigm of automation and self-adaptation. Experimental evaluation across four representative HPC applications demonstrates that the system achieves expert-level linear scalability, substantially reduces job completion time, and yields actionable best practices for collaborative agent design in HPC contexts.
This work addresses the challenges of workflow task composition in high-throughput, petabyte-scale data processing environments, where resource heterogeneity and execution overhead significantly impact performance. The authors propose a hybrid task composition strategy that dynamically balances task independence against execution grouping, formulated within a multi-objective optimization framework to achieve Pareto-optimal trade-offs among throughput, I/O cost, and CPU efficiency. Leveraging workflow DAG modeling and high-dimensional parameter space simulation, the approach enables policy-driven automated synthesis of workflows. Experimental results demonstrate that the proposed strategy achieves up to a 3.8× improvement in throughput and reduces network overhead by as much as 14.9× compared to baseline methods, offering a scalable workflow synthesis framework for extreme-scale scientific computing.
This work addresses the orchestration bottlenecks faced by ultra-large-scale Sim-AI workflows on leadership-class supercomputers, which arise from task heterogeneity and extreme ensemble sizes. To overcome these challenges, the authors propose EnsembleLauncher, a recursively hierarchical and fully decentralized workflow orchestrator that introduces a decentralized control plane and a programmable scheduling policy interface, thereby surpassing conventional tools in both scalability and scheduling flexibility. Experiments on the Aurora supercomputer demonstrate that EnsembleLauncher can efficiently schedule system-wide resources to support up to 8 million serial tasks, achieving more than a fourfold performance improvement over state-of-the-art alternatives. Furthermore, it significantly enhances resource utilization for workloads with high task variance and active learning pipelines.
This work addresses the challenge of automated resource discovery and scheduling across heterogeneous, multi-institutional computing environments—including cloud, edge, and high-performance computing (HPC) infrastructures—in agent-based systems. The authors propose a hierarchical dynamic agent architecture in which secretary agents concurrently and asynchronously perform resource probing, negotiation, and task dispatching. By integrating an asynchronous negotiation protocol, agent-driven resource categorization, and a dynamic scheduling algorithm, the framework enables highly scalable, cross-infrastructure automation. Evaluated on a testbed comprising 51 real and simulated resource providers, the approach achieved a negotiation accuracy of 87.71% over 19,973 negotiation rounds and 6,952 task selections, with task selection costs comparable to conventional strategies, thereby significantly enhancing scheduling efficiency and adaptability.