Score
Designs, implements, and operates systems that coordinate, schedule, and manage the execution of tasks, services, containers, agents, pipelines, experiments, models, and resources across clusters and environments. Builds or evaluates orchestration frameworks and tooling for deployment and pipeline workflows (ETL, training, inference), multi‑agent and service coordination, API/tool integrations, and resource/task allocation including Kubernetes/container and cluster orchestration.
Existing Kubernetes schedulers struggle to simultaneously optimize user-defined QoS objectives—such as energy efficiency, cost, and global performance—while lacking automated, declarative orchestration capabilities across heterogeneous cloud-fog-edge clusters. To address this, we propose the first QoS-aware federated orchestration system. Our approach employs a lightweight, Raft-replicated resource agent architecture tightly coupled with a centralized knowledge repository, enabling, for the first time, automatic translation of user-specified YAML-declared multi-dimensional QoS constraints (e.g., latency, energy consumption, cost) into microservice placement and dynamic migration policies. The system integrates Istio service mesh and federated cluster management to support policy-driven scheduling, QoS-compliant rescheduling, and zero-touch failover. Evaluated on a nine-cluster testbed, our system demonstrates both effectiveness and scalability in meeting diverse, cross-layer QoS requirements.
This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.
This work proposes CODECO, a framework designed to address the challenges of traditional centralized Kubernetes in federated edge environments, where heterogeneous infrastructure, device mobility, and multi-provider collaboration are prevalent. CODECO enables edge autonomy while preserving global consistency through co-orchestration of data, computation, and networking. It integrates a semantic application model, a partitioned federation mechanism, AI-driven scheduling decisions, and a hybrid governance model. Built upon an extended Kubernetes architecture, CODECO supports context-aware microservice deployment and adaptive management. The framework’s efficacy in orchestrating applications across complex federated edge-cloud scenarios is validated through a reproducible experimental platform, demonstrating its capability to efficiently manage dynamic and heterogeneous edge environments.
This work addresses the limitations of Kubernetes’ default scheduler, which often leads to resource fragmentation and suboptimal utilization due to its local decision-making nature, while existing global scheduling approaches struggle with practical deployment in production clusters. The paper proposes OPSche, the first open-source plugin that collaboratively operates alongside the default scheduler by leveraging the Kubernetes scheduling framework to seamlessly integrate globally optimized schedules generated by external solvers through atomic validation and coordination hooks. OPSche supports three trigger modes—scheduling failure, periodic invocation, and queue stabilization—along with their blocking variants, thereby balancing scheduling quality, latency, and interference without replacing the native scheduler. Experimental results demonstrate that OPSche improves resource utilization by up to 3.0% across diverse cluster configurations and reduces scheduling latency by over one second.
To address low application deployment/reconfiguration efficiency and suboptimal resource utilization in Cooperative Intelligent Transportation Systems (C-ITS) under dynamic environments and heterogeneous, multi-source requirements, this paper proposes a requirement-driven cloud-native application management approach. The method innovatively integrates Kubernetes container orchestration with the ROS 2 real-time communication framework, establishing an automated management architecture that supports on-demand microservice deployment, dynamic reconfiguration, and elastic scaling—enabling plug-and-play integration and closed-loop response of external auxiliary services. The framework is deeply optimized for edge–cloud collaboration in C-ITS. Evaluated in a collective environmental perception use case, it reduces computational resource consumption by 32% and network traffic by 41%. The prototype system is open-sourced.
This study addresses the challenge of automating workflows in complex industries—such as logistics, healthcare, and construction—where processes are fragmented across heterogeneous tools and involve multi-party collaboration. The work proposes orchestration as a core abstraction to enable effective automation by dynamically coordinating multi-step tasks, enforcing domain-specific constraints, managing human approvals, and integrating legacy systems. It introduces the novel concept of “orchestration bottlenecks” and develops a theoretical framework that unifies multi-agent systems, workflow modeling, constraint reasoning, and human–AI collaboration, while exposing critical gaps in current multi-agent approaches at the orchestration level. Based on distinct sources of operational friction across domains, the paper advocates for targeted architectural safeguards—such as constraint enforcement or explainability—and phased implementation strategies to provide actionable pathways for automation in complex operational environments.
This work addresses the challenges of scheduling scientific workflows in multi-cluster environments, where dynamic task transformation, heterogeneous resource matching, and reliable job dispatch are critical. The authors propose an intelligent agent framework powered by large language models that leverages multidimensional prompt engineering to parse natural-language instructions and automatically generate job specifications. Integrated with a metadata-driven dynamic transformation mechanism, the framework enables end-to-end automation of cross-cluster scheduling. Experimental results demonstrate a 97.9% scheduling success rate across 432 runs; incorporating descriptive metadata increased the successful execution rate of 220 jobs from 48% to 87%, while five representative applications achieved up to a 3.3× performance speedup, effectively resolving architectural mismatches and enabling efficient, reliable task distribution.
This work proposes the first large language model (LLM)-driven agent framework for autonomous, end-to-end management of high-performance computing (HPC) applications in cloud environments. Addressing the heavy reliance on manual intervention and the lack of intelligent decision-making in traditional HPC cloud deployment, the framework enables automated multi-platform container construction, Kubernetes-based orchestration, cross-instance performance optimization, and adaptive elastic scaling policy generation. By integrating LLM-powered agents into HPC cloud workflow orchestration, this study establishes a novel paradigm of automation and self-adaptation. Experimental evaluation across four representative HPC applications demonstrates that the system achieves expert-level linear scalability, substantially reduces job completion time, and yields actionable best practices for collaborative agent design in HPC contexts.
This study addresses the challenges of high latency, unstable concurrency, and security risks faced by large language model (LLM) agents in automating asset lifecycle management within Industry 4.0. The authors propose a Plan-then-Execute architecture that generates verifiable workflow graphs and integrates a topology-aware parallel scheduling mechanism to enable controlled inference overlap while ensuring functional correctness and security. Key technical contributions include topological-sort-based multi-agent scheduling, structured context pruning, dependency-aware concurrency control, and graceful degradation under fault injection. Evaluated on the AssetOpsBench benchmark, the system reduces median end-to-end latency by 1.6× (up to 1.8× for highly parallel tasks) and cuts inference overhead by approximately 30% through context pruning, all while maintaining stable task completion rates and output quality.
The recent convergence of edge computing, serverless execution, and Kubernetes (K8s) based container orchestration has enabled the processing of application workflows close to data sources. While effective within a single edge cluster, existing schemes do not generalize to federated multi edge environments, where multiple workflows execute concurrently under strict end to end (E2E) deadline constraints. This paper introduces ClusterLess, a deadline aware serverless workflow orchestration method for federated multi edge K8s clusters. ClusterLess manages the E2E lifecycle of workflow execution, including dependency analysis, execution mode selection, and resource aware placement. To this end, it integrates structured intra cluster orchestration with a leader selected, super master driven intercluster coordination layer, determining where and how each workflow function should be executed across the federated edge clusters. We implement ClusterLess using OpenFaaS as the serverless execution substrate and Argo for workflow management, and deploy it on a realistic testbed of six edge clusters comprising 64 heterogeneous edge nodes. Experimental results with concurrent serverless workflows, spanning 18 workload configurations across different input sizes and deadline classes, show that ClusterLess reduces workflow completion time by up to 40 %, increases deadline satisfaction from below 50 % to over 90 %, and confines deadline violations to single digit seconds compared to four baseline methods.