hybrid cluster architecture design

Designs the structure and component layout of computing clusters that combine heterogeneous resources or deployment models (for example, mixes of on‑premises and cloud nodes, CPUs and accelerators, or virtualized and bare‑metal systems), specifying node types, interconnects, storage tiers, placement and scheduling policies, and fault‑tolerance mechanisms to meet performance, scalability, and cost goals. Involves building deployment topologies, resource management and workload placement strategies, and analyzing trade‑offs among latency, throughput, data locality, reliability, and operational cost.

hybridclusterarchitecturedesign

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.32
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of resource allocation in geographically distributed and heterogeneous continuum computing infrastructures, where combinatorial explosion and limited generalization hinder effective deployment. To tackle this, the study introduces, for the first time, the pricing structures commonly found in Software-as-a-Service (SaaS) ecosystems into the resource allocation problem, formulating a unified, price-based representation of the configuration space. The authors propose PRIME, a pricing-aware analysis engine that efficiently searches for cost-optimal deployment configurations satisfying both functional and non-functional constraints. Leveraging synthetic infrastructure topologies and workload generation techniques, the project constructs a comprehensive dataset comprising 9,600 diverse scenarios, demonstrating that the proposed approach achieves both scalability and computational efficiency in complex, heterogeneous environments.

computing continuumconfiguration spaceheterogeneous infrastructure

Scientific computing in heterogeneous environments faces significant challenges in simultaneously achieving high performance, cost efficiency, scalability, and accessibility. This work proposes a hybrid cloud architecture tailored for scientific computing that integrates grid and cloud platforms—such as SLURM, OpenPBS, OpenStack, and Kubernetes—with workflow systems including Nextflow, Snakemake, and Common Workflow Language (CWL). By leveraging federated computing, multi-cloud orchestration, and a unified governance framework, the architecture enables seamless cross-platform resource scheduling and task coordination. The approach substantially enhances infrastructure interoperability and sustainability, with validation in life sciences demonstrating its practical efficacy. It has already facilitated integration and large-scale adoption within the ELIXIR and European Open Science Cloud (EOSC) ecosystems.

computing infrastructurehybrid cloudinteroperability

An Analysis of HPC and Edge Architectures in the Cloud

Aug 02, 2025
SS
Steven Santillan
🏛️ Escuela Superior Politécnica del Litoral | ESPOL

This study addresses the practical disparities and co-evolution between high-performance computing (HPC) and edge computing architectures within the cloud continuum. It presents the first large-scale empirical analysis based on 396 real-world, production-grade AWS architectures. Methodologically, we propose a multidimensional, data-driven framework encompassing service topology identification, storage type classification, architectural complexity quantification, and ML service integration statistics. Results reveal systematic differences—and complementary patterns—between HPC and edge architectures across four dimensions: core service composition (e.g., EC2 versus Greengrass/Lambda), storage design paradigms (parallel file systems versus distributed lightweight caches), complexity distributions, and ML embedding strategies. This work delivers the first industry-scale architectural benchmark for the cloud continuum, providing empirically grounded insights and methodological foundations for cross-domain architecture design, resource optimization, and cloud-native convergence of HPC and edge computing.

Analyze HPC and edge architectures in AWS cloud deploymentsAssess architectural complexity and machine learning services usageInvestigate AWS services prevalence and storage systems used

To address node overload, high operational costs, and poor system stability caused by dynamic heterogeneous resource scheduling in cloud computing, this paper proposes an intelligent load-balancing framework. The method constructs a high-fidelity simulation environment and an abstracted multi-resource model, introducing for the first time a joint resource utilization metric that incorporates VM migration overhead. It establishes a novel three-category taxonomy for schedulers, derives an empirically grounded formula for estimating VM migration traffic, and comparatively evaluates two emerging paradigms: centralized metaheuristic and distributed multi-agent scheduling. Built upon real-world Google cluster traces, the framework integrates live VM migration and realistic workload simulation. Experimental validation on the University of Westminster’s HPC cluster demonstrates a 23.6% improvement in resource utilization, a 31.4% reduction in task latency, and a 27.9% decrease in network migration overhead—significantly enhancing system stability and cost-efficiency.

Designing dynamic task allocation to prevent cloud node overloadDeveloping strategies to maintain system stability at minimal costProposing centralized and decentralized approaches for resource management

Resource Management Schemes for Cloud-Native Platforms with Computing Containers of Docker and Kubernetes

Oct 20, 2020
YM
Ying Mao
🏛️ Fordham University | Dublin City University | Wageningen University

This study addresses prolonged task completion times, low resource utilization, and high resource release latency in Docker/Kubernetes containers on cloud-native platforms running compute-intensive workloads (e.g., big data and deep learning). We systematically evaluate the performance impact of diverse resource scheduling strategies through system-level monitoring—leveraging cgroups and metrics-server—and multi-workload stress testing. For the first time, we empirically quantify how key resource configurations significantly affect task completion time (±79.4% variation) and resource release latency (+116.7% degradation). Based on these findings, we propose an evidence-driven configuration optimization paradigm that reduces maximum task completion time by up to 79.4% and precisely identifies configuration bottlenecks responsible for latency. Our results provide reproducible, transferable empirical foundations for resource management tuning and deployment decisions in cloud-native environments.

Analyzing system overhead and resource usage in cloud-native environmentsEvaluating performance of big data and deep learning applicationsInvestigating resource management schemes for Docker and Kubernetes platforms

Latest Papers

What's happening recently
View more

The deployment of high-density AI accelerators renders traditional data center power delivery hierarchies inefficient in utilizing provisioned power, leading to resource waste and stranded capacity. This work presents the first integrated evaluation framework that jointly models GPU, compute, and storage placement, leveraging real-world Azure traces of workload arrivals, overbooking patterns, and hardware retirement schedules to co-optimize power, performance, and cost. Innovatively incorporating multi-resource stranding into power infrastructure design, the study proposes “deployable capacity”—the actual computational capacity that can be effectively powered and utilized—as a more meaningful planning objective than conventional “nameplate power.” It further quantifies the impact of high-density AI systems on deployable capacity, effective capital expenditure, and delivered performance, offering critical insights for rethinking data center power architectures in the AI era.

AI acceleratorsdatacenterpower delivery

This work addresses the challenges of constructing high-density distributed space-based data centers in low Earth orbit (LEO) by proposing two parametric satellite constellation architectures—planar and three-dimensional—that optimize geometric layouts under constraints including minimum inter-satellite distance, unobstructed solar power access, and stable inter-satellite links. The design maps a VL2-inspired Clos network topology onto feasible inter-satellite links. Leveraging orbital dynamics modeling, numerical analysis, and integer optimization, the approach achieves densest packing in the planar configuration and enables the 3D architecture to scale satellite count proportionally to $(R_{\max}/R_{\min})^3$. Experimental results demonstrate that both architectures provide sufficient persistent, obstruction-free links to replicate terrestrial data center switching fabrics, while quantifying the trade-off between per-satellite link capacity and the number of dedicated switching satellites.

distributed space-based datacentersinter-satellite linksLEO constellations

This work addresses the lack of systematic, reproducible, and maintainable testing methodologies in existing dynamic resource management libraries. We propose an automated validation framework tailored for high-performance computing (HPC) environments, which introduces a novel multi-level testing taxonomy encompassing both functional and non-functional requirements. Built upon an MPI-based scalable library testing methodology, the framework supports core primitives of dynamic resource management systems—such as initialization, readiness checks, and reconfiguration—and integrates containerized virtual clusters with continuous integration (CI) ecosystems. Experimental evaluation demonstrates that our approach significantly improves early fault detection rates, reduces maintenance overhead caused by evolving dependencies, and is readily generalizable to other systems exhibiting similar variability mechanisms.

dynamic resource managementHPClibrary correctness