Score
Designs and builds automated, reproducible systems that coordinate the deployment and lifecycle of containerized services, clusters, networks, pipelines, and toolchains across cloud environments (including Kubernetes), enabling rollout, scaling, sandboxed experiments, and real‑time operation. Analyzes and implements vendor‑agnostic, multi‑vendor orchestration and control layers to provision resources, manage service and pipeline interactions, and enforce predictable deployment and teardown workflows.
This study addresses the challenge of automating the deployment of multi-service containerized applications in heterogeneous edge-cloud environments, where minimizing manual intervention while ensuring performance guarantees remains difficult. The authors evaluate and validate the open-source CODECO toolkit’s orchestration capabilities across diverse hardware platforms—including ARM, AMD, and Raspberry Pi—and lightweight Kubernetes distributions such as k3s. Experimental results demonstrate that, compared to standard Kubernetes workflows, CODECO significantly reduces human intervention during deployment while maintaining competitive deployment efficiency, runtime performance, and manageable resource overhead. These findings highlight CODECO’s strong compatibility and its potential to lower operational complexity in edge-cloud collaborative deployments.
This work proposes CODECO, a framework designed to address the challenges of traditional centralized Kubernetes in federated edge environments, where heterogeneous infrastructure, device mobility, and multi-provider collaboration are prevalent. CODECO enables edge autonomy while preserving global consistency through co-orchestration of data, computation, and networking. It integrates a semantic application model, a partitioned federation mechanism, AI-driven scheduling decisions, and a hybrid governance model. Built upon an extended Kubernetes architecture, CODECO supports context-aware microservice deployment and adaptive management. The framework’s efficacy in orchestrating applications across complex federated edge-cloud scenarios is validated through a reproducible experimental platform, demonstrating its capability to efficiently manage dynamic and heterogeneous edge environments.
Existing Kubernetes schedulers struggle to simultaneously optimize user-defined QoS objectives—such as energy efficiency, cost, and global performance—while lacking automated, declarative orchestration capabilities across heterogeneous cloud-fog-edge clusters. To address this, we propose the first QoS-aware federated orchestration system. Our approach employs a lightweight, Raft-replicated resource agent architecture tightly coupled with a centralized knowledge repository, enabling, for the first time, automatic translation of user-specified YAML-declared multi-dimensional QoS constraints (e.g., latency, energy consumption, cost) into microservice placement and dynamic migration policies. The system integrates Istio service mesh and federated cluster management to support policy-driven scheduling, QoS-compliant rescheduling, and zero-touch failover. Evaluated on a nine-cluster testbed, our system demonstrates both effectiveness and scalability in meeting diverse, cross-layer QoS requirements.
Research on containerization in multi-cloud environments remains fragmented, lacking a systematic, up-to-date synthesis. Method: We conduct a Systematic Mapping Study (SMS) spanning 2013–2024, analyzing 121 high-quality publications through bibliometric analysis, thematic coding, and ISO/IEC 25010 quality attribute modeling. Contribution/Results: We propose the first four-level classification framework—“Theme–Strategy–Quality Attribute–Tactic”—identifying four core research themes, 98 implementation strategies, 10 critical quality attributes, and 47 corresponding architectural tactics. Innovatively, we introduce a two-dimensional challenge-solution taxonomy organized along Security, Automation, Deployment, and Monitoring dimensions. This yields the first structured, reusable landscape of multi-cloud containerization, bridging theoretical research and industrial practice by supporting architecture design and technology selection—thereby addressing a longstanding gap in systematic knowledge integration for this domain.
To address low application deployment/reconfiguration efficiency and suboptimal resource utilization in Cooperative Intelligent Transportation Systems (C-ITS) under dynamic environments and heterogeneous, multi-source requirements, this paper proposes a requirement-driven cloud-native application management approach. The method innovatively integrates Kubernetes container orchestration with the ROS 2 real-time communication framework, establishing an automated management architecture that supports on-demand microservice deployment, dynamic reconfiguration, and elastic scaling—enabling plug-and-play integration and closed-loop response of external auxiliary services. The framework is deeply optimized for edge–cloud collaboration in C-ITS. Evaluated in a collective environmental perception use case, it reduces computational resource consumption by 32% and network traffic by 41%. The prototype system is open-sourced.
This work addresses the limitations of traditional single-tenant batch systems in meeting the demands of AI training, secure computation on sensitive data, and mixed workloads requiring flexibility and reproducibility. It presents the first deployment of a multi-tenant Kubernetes infrastructure on the HPE Cray EX supercomputer (Isambard-AI), integrating trusted research environments with distributed AI model hosting services. The proposed architecture leverages KubeRay, Ray, vLLM, and HPE Slingshot interconnects to deliver a sandboxed, persistent platform. By extending Kubernetes beyond conventional cloud environments into high-performance bare-metal systems, this study demonstrates the feasibility of scalable “Kubernetes-as-a-Service” for national-scale AI infrastructure. It further identifies key implementation challenges and outlines an evolutionary pathway, offering a co-design paradigm for multi-tenant confidential computing in domains such as healthcare.
The recent convergence of edge computing, serverless execution, and Kubernetes (K8s) based container orchestration has enabled the processing of application workflows close to data sources. While effective within a single edge cluster, existing schemes do not generalize to federated multi edge environments, where multiple workflows execute concurrently under strict end to end (E2E) deadline constraints. This paper introduces ClusterLess, a deadline aware serverless workflow orchestration method for federated multi edge K8s clusters. ClusterLess manages the E2E lifecycle of workflow execution, including dependency analysis, execution mode selection, and resource aware placement. To this end, it integrates structured intra cluster orchestration with a leader selected, super master driven intercluster coordination layer, determining where and how each workflow function should be executed across the federated edge clusters. We implement ClusterLess using OpenFaaS as the serverless execution substrate and Argo for workflow management, and deploy it on a realistic testbed of six edge clusters comprising 64 heterogeneous edge nodes. Experimental results with concurrent serverless workflows, spanning 18 workload configurations across different input sizes and deadline classes, show that ClusterLess reduces workflow completion time by up to 40 %, increases deadline satisfaction from below 50 % to over 90 %, and confines deadline violations to single digit seconds compared to four baseline methods.
This work proposes the first large language model (LLM)-driven agent framework for autonomous, end-to-end management of high-performance computing (HPC) applications in cloud environments. Addressing the heavy reliance on manual intervention and the lack of intelligent decision-making in traditional HPC cloud deployment, the framework enables automated multi-platform container construction, Kubernetes-based orchestration, cross-instance performance optimization, and adaptive elastic scaling policy generation. By integrating LLM-powered agents into HPC cloud workflow orchestration, this study establishes a novel paradigm of automation and self-adaptation. Experimental evaluation across four representative HPC applications demonstrates that the system achieves expert-level linear scalability, substantially reduces job completion time, and yields actionable best practices for collaborative agent design in HPC contexts.
This work addresses the growing challenge in modern scientific research where complex infrastructure management, authentication, and deployment processes divert focus from core scientific discovery. To this end, the authors propose Sci-Orchestra, a cloud-native, Kubernetes-based hierarchical orchestration framework that automates experimental workflows through API-driven mechanisms, enabling secure authentication, resource scheduling, and scalable deployment across heterogeneous high-performance computing environments. Sci-Orchestra introduces an autonomous service marketplace to facilitate cross-institutional collaboration and adopts a “black-box” interoperability model that promotes integration of academic, industrial, and research tools while safeguarding intellectual property. By significantly lowering technical barriers, the framework accelerates the transition of research prototypes into production-grade applications, thereby advancing the emerging paradigm of Science as a Service (SciaaS).