Score
Designs, builds, and analyzes architectures and operational platforms that integrate on‑premises, private, and public cloud resources into a cohesive hybrid environment, addressing networking, identity, security, workload placement, data synchronization, and orchestration across boundaries. This includes defining deployment and management patterns, pipelines, and lifecycle controls for workloads—such as AI models and their data—when those workloads must span hybrid infrastructures.
Scientific computing in heterogeneous environments faces significant challenges in simultaneously achieving high performance, cost efficiency, scalability, and accessibility. This work proposes a hybrid cloud architecture tailored for scientific computing that integrates grid and cloud platforms—such as SLURM, OpenPBS, OpenStack, and Kubernetes—with workflow systems including Nextflow, Snakemake, and Common Workflow Language (CWL). By leveraging federated computing, multi-cloud orchestration, and a unified governance framework, the architecture enables seamless cross-platform resource scheduling and task coordination. The approach substantially enhances infrastructure interoperability and sustainability, with validation in life sciences demonstrating its practical efficacy. It has already facilitated integration and large-scale adoption within the ELIXIR and European Open Science Cloud (EOSC) ecosystems.
This study addresses the joint optimization of performance, cost, and regulatory compliance in hybrid cloud deployments—challenged by resource allocation imbalance, heterogeneous multi-cloud pricing models, and inadequate protection of sensitive data. We propose the first unified security–cost co-optimization framework integrating zero-trust architecture, dynamic encryption policies, and a cross-cloud policy engine. Leveraging Policy-as-Code (PaC), the framework enables coordinated decision-making for security enforcement and resource scheduling. It is operationalized via native integration with AWS and Azure APIs and supports hybrid cloud orchestration. Evaluated in a real-world AWS+Azure environment, our approach achieves a 40% improvement in security incident response latency, reduces cross-cloud resource misallocation by 52%, and attains 100% compliance audit pass rate.
This study addresses the coordination challenges arising from independent control across cloud, high-performance computing (HPC), and edge AI infrastructures. Conceptualizing the AI platform as a "system of systems," this work proposes an architectural paradigm characterized by usage fusion and federated control. Methodologically, it adopts a systems engineering framework that achieves cross-domain coordination through interface contracts while preserving native control planes. The approach incorporates boundary testing, responsibility models, and seven integration facets, leveraging interface mapping, policy contexts, and operational evidence to guide integration design. The primary contribution lies in establishing a unified framework for evaluating interoperability, governance capabilities, and fault isolation, thereby delineating clear directions for future research.
This study addresses the practical disparities and co-evolution between high-performance computing (HPC) and edge computing architectures within the cloud continuum. It presents the first large-scale empirical analysis based on 396 real-world, production-grade AWS architectures. Methodologically, we propose a multidimensional, data-driven framework encompassing service topology identification, storage type classification, architectural complexity quantification, and ML service integration statistics. Results reveal systematic differences—and complementary patterns—between HPC and edge architectures across four dimensions: core service composition (e.g., EC2 versus Greengrass/Lambda), storage design paradigms (parallel file systems versus distributed lightweight caches), complexity distributions, and ML embedding strategies. This work delivers the first industry-scale architectural benchmark for the cloud continuum, providing empirically grounded insights and methodological foundations for cross-domain architecture design, resource optimization, and cloud-native convergence of HPC and edge computing.
To address mounting challenges—including usability, manageability, energy efficiency, cost, and scalability—posed by increasingly complex AI workloads, this project proposes a full-stack, co-designed hybrid cloud rearchitecting framework. Methodologically, it introduces four key innovations: (1) the LLM-as-Abstraction (LLMaaA) paradigm; (2) AI-agent-driven cross-layer automation; (3) quantum-classical hybrid workflows; and (4) a physics-enhanced scientific AI agent framework. By integrating generative AI and multi-agent systems, the framework establishes a unified control plane, a composable adaptive architecture, and an edge-cloud collaborative programming model. Evaluated on high-impact applications—including materials discovery and climate modeling—the platform achieves a 32% improvement in task completion rate and a 27% reduction in energy consumption, thereby significantly enhancing system sustainability, security, and operational efficiency.
This study addresses the interoperability and migration challenges enterprises face when deploying workloads across AWS and Alibaba Cloud. Through a systematic comparison of architectural designs, service offerings, and operational policies between the two platforms, the research conducts an exploratory case study on migrating IoT workloads using both native and open-source Infrastructure-as-Code (IaC) tools. It reveals critical technical trade-offs inherent in cross-cloud co-deployment for the first time, distills best practices for secure, resilient, and vendor-lock-in-mitigated multicloud deployments, and proposes a multicloud interoperability framework tailored for global enterprises. The findings offer methodological support for empirically grounded multicloud strategies.
This study addresses the infrastructure complexity of cloud-edge-end协同 architectures, which has emerged as a major bottleneck hindering developer productivity and innovation. Through 101 semi-structured interviews across 86 organizations, this work empirically identifies deployment complexity and onboarding difficulty as core challenges. It proposes four architectural directions to mitigate these issues: Object-as-a-Service (unified object abstraction), internal developer platforms, declarative AI/ML pipelines, and lightweight edge runtimes. Findings indicate that high-level abstractions and automation significantly enhance developer experience—outweighing the impact of execution performance optimizations—and thereby establish a new paradigm for platform engineering and distributed system design.
This study addresses the challenges faced by edge-cloud-native applications in cross-industry adoption, including fragmented toolchains, steep learning curves, and inconsistent performance across hybrid environments. Through in-depth interviews with practitioners from multiple sectors, the work offers the first systematic insight into the real-world pain points experienced by non-technical teams during digital transformation. It reveals that such teams prioritize productivity, service quality, and usability over cost alone. Grounded in qualitative analysis, the research identifies key platform design requirements centered on developer-friendliness, end-to-end lifecycle simplification, and SLA-aware orchestration, with a focus on distributed network computing, hybrid cloud management, and service-level agreement (SLA) assurance. These findings provide a practice-oriented roadmap for the evolution of converged cloud-network infrastructures.
This work addresses the challenge that existing experimental environments for distributed Cyber-Physical Systems (CPS) struggle to support reproducible, observable, and controllable integration of heterogeneous edge, fog, and cloud resources. To bridge this gap, the paper proposes a generic cloud continuum experimentation architecture grounded in the SLICES blueprint, featuring a two-layer reference model that decouples infrastructure from application logic. CPS workflows are structured along an edge–fog–cloud continuum, with deployment location, timing, and data provenance treated as core experimental dimensions. The architecture integrates virtualized and physical edge nodes, digital twin coordination, time-windowed control, and combined stream processing with cloud-side aggregation analytics, enabling multi-domain CPS applications to share programmable infrastructure and flexibly deploy and compare control and monitoring strategies. Validation through 40 systematic experiments across geographically distributed deployments—spanning renewable energy community management and AirWatch monitoring use cases—demonstrates the framework’s effectiveness and generality in hybrid physical-virtual settings.
This work addresses the challenges of resource utilization and operational efficiency in microservice architectures by proposing a performance-metric-driven automated framework that intelligently determines the optimal deployment strategy for individual microservices between Infrastructure-as-a-Service (IaaS) and Function-as-a-Service (FaaS). By analyzing intrinsic microservice characteristics, the framework enables a scalable and reproducible migration from conventional IaaS deployments to a hybrid IaaS+FaaS model. Experimental evaluation on two real-world applications demonstrates that the approach accurately identifies microservices well-suited for serverless execution, significantly improving both deployment efficiency and resource utilization. Furthermore, the study clarifies the respective applicability boundaries and advantages of different cloud service models, offering practical guidance for architecture design in heterogeneous cloud environments.