cloud platforms

Designs, builds, deploys, and operates internet-delivered compute, storage, networking, and platform services (public, private, or hybrid); defines architecture, resource provisioning, security controls, scaling and resilience, monitoring, and cost management. Implements and automates cloud infrastructure and platform components using provider APIs and infrastructure-as-code, and analyzes performance, reliability, and operational telemetry to optimize workloads.

cloudplatforms

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
1.19
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$192K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Cloud Infrastructure Management in the Age of AI Agents

Jun 13, 2025
ZY
Zhenning Yang
🏛️ University of Michigan | UC Berkeley | Andreessen Horowitz

This study addresses the high manual overhead faced by DevOps teams in managing multi-interface cloud infrastructures. We propose and systematically evaluate an LLM-driven AI agent framework for automation. Methodologically, the agent unifies heterogeneous interfaces—including SDKs, CLIs, Infrastructure-as-Code (IaC) tools, and web portals—to support core tasks such as configuration deployment, monitoring/alerting, and incident remediation. Key contributions include: (1) the first evaluation framework specifically designed for AI agents in cloud infrastructure management; (2) identification and systematic mitigation of three critical bottlenecks—interface semantic gaps, action execution reliability, and security constraint compliance; and (3) domain-specific optimization strategies validated in real-world deployments, demonstrating both task feasibility and cross-scenario generalizability. Our work establishes a reusable methodology and empirical benchmark for AI-native cloud operations.

Automating cloud infrastructure management using AI agentsEvaluating AI agents across diverse cloud interfacesIdentifying challenges in AI-driven cloud management solutions

This study addresses the infrastructure complexity of cloud-edge-end协同 architectures, which has emerged as a major bottleneck hindering developer productivity and innovation. Through 101 semi-structured interviews across 86 organizations, this work empirically identifies deployment complexity and onboarding difficulty as core challenges. It proposes four architectural directions to mitigate these issues: Object-as-a-Service (unified object abstraction), internal developer platforms, declarative AI/ML pipelines, and lightweight edge runtimes. Findings indicate that high-level abstractions and automation significantly enhance developer experience—outweighing the impact of execution performance optimizations—and thereby establish a new paradigm for platform engineering and distributed system design.

cloud-edge infrastructuredeveloper productivitydistributed computing

An Analysis of HPC and Edge Architectures in the Cloud

Aug 02, 2025
SS
Steven Santillan
🏛️ Escuela Superior Politécnica del Litoral | ESPOL

This study addresses the practical disparities and co-evolution between high-performance computing (HPC) and edge computing architectures within the cloud continuum. It presents the first large-scale empirical analysis based on 396 real-world, production-grade AWS architectures. Methodologically, we propose a multidimensional, data-driven framework encompassing service topology identification, storage type classification, architectural complexity quantification, and ML service integration statistics. Results reveal systematic differences—and complementary patterns—between HPC and edge architectures across four dimensions: core service composition (e.g., EC2 versus Greengrass/Lambda), storage design paradigms (parallel file systems versus distributed lightweight caches), complexity distributions, and ML embedding strategies. This work delivers the first industry-scale architectural benchmark for the cloud continuum, providing empirically grounded insights and methodological foundations for cross-domain architecture design, resource optimization, and cloud-native convergence of HPC and edge computing.

Analyze HPC and edge architectures in AWS cloud deploymentsAssess architectural complexity and machine learning services usageInvestigate AWS services prevalence and storage systems used

Over-the-Top Resource Broker System for Split Computing: An Approach to Distribute Cloud Computing Infrastructure

Aug 11, 2025
IF
Ingo Friese
🏛️ Deutsche Telekom AG | Intelligent Networks | German Research Center for Artificial Intelligence (DFKI)

To address the challenges of cross-operator heterogeneous infrastructure coordination, complex service deployment, and insufficient performance guarantees in 6G networks, this paper proposes an Over-the-Top Resource Broker—a network-computing co-designed, cross-layer resource brokerage architecture. The architecture leverages split computing, multi-level resource virtualization, service-level agreement (SLA) negotiation protocols, and distributed state management to enable unified abstraction and fine-grained, location- and service-aware scheduling across dynamic cloud-edge-end nodes. It represents the first extension of the resource broker paradigm to native 6G scenarios, supporting multi-provider resource sharing and performance-guaranteed, on-demand service provisioning. Prototype evaluation demonstrates that the system significantly reduces cross-domain service integration complexity while improving resource utilization and scheduling flexibility.

Abstract complexities of multiple infrastructure providers for seamless deploymentDistribute cloud computing infrastructure for 6G split computingEnsure uniform interface across diverse processing node characteristics

Resource Management Schemes for Cloud-Native Platforms with Computing Containers of Docker and Kubernetes

Oct 20, 2020
YM
Ying Mao
🏛️ Fordham University | Dublin City University | Wageningen University

This study addresses prolonged task completion times, low resource utilization, and high resource release latency in Docker/Kubernetes containers on cloud-native platforms running compute-intensive workloads (e.g., big data and deep learning). We systematically evaluate the performance impact of diverse resource scheduling strategies through system-level monitoring—leveraging cgroups and metrics-server—and multi-workload stress testing. For the first time, we empirically quantify how key resource configurations significantly affect task completion time (±79.4% variation) and resource release latency (+116.7% degradation). Based on these findings, we propose an evidence-driven configuration optimization paradigm that reduces maximum task completion time by up to 79.4% and precisely identifies configuration bottlenecks responsible for latency. Our results provide reproducible, transferable empirical foundations for resource management tuning and deployment decisions in cloud-native environments.

Analyzing system overhead and resource usage in cloud-native environmentsEvaluating performance of big data and deep learning applicationsInvestigating resource management schemes for Docker and Kubernetes platforms

Latest Papers

What's happening recently
View more

This work addresses the challenge of resource allocation in geographically distributed and heterogeneous continuum computing infrastructures, where combinatorial explosion and limited generalization hinder effective deployment. To tackle this, the study introduces, for the first time, the pricing structures commonly found in Software-as-a-Service (SaaS) ecosystems into the resource allocation problem, formulating a unified, price-based representation of the configuration space. The authors propose PRIME, a pricing-aware analysis engine that efficiently searches for cost-optimal deployment configurations satisfying both functional and non-functional constraints. Leveraging synthetic infrastructure topologies and workload generation techniques, the project constructs a comprehensive dataset comprising 9,600 diverse scenarios, demonstrating that the proposed approach achieves both scalability and computational efficiency in complex, heterogeneous environments.

computing continuumconfiguration spaceheterogeneous infrastructure

This work addresses the challenges of resource utilization and operational efficiency in microservice architectures by proposing a performance-metric-driven automated framework that intelligently determines the optimal deployment strategy for individual microservices between Infrastructure-as-a-Service (IaaS) and Function-as-a-Service (FaaS). By analyzing intrinsic microservice characteristics, the framework enables a scalable and reproducible migration from conventional IaaS deployments to a hybrid IaaS+FaaS model. Experimental evaluation on two real-world applications demonstrates that the approach accurately identifies microservices well-suited for serverless execution, significantly improving both deployment efficiency and resource utilization. Furthermore, the study clarifies the respective applicability boundaries and advantages of different cloud service models, offering practical guidance for architecture design in heterogeneous cloud environments.

Cloud DeploymentFaaSIaaS

This work addresses the challenge that existing experimental environments for distributed Cyber-Physical Systems (CPS) struggle to support reproducible, observable, and controllable integration of heterogeneous edge, fog, and cloud resources. To bridge this gap, the paper proposes a generic cloud continuum experimentation architecture grounded in the SLICES blueprint, featuring a two-layer reference model that decouples infrastructure from application logic. CPS workflows are structured along an edge–fog–cloud continuum, with deployment location, timing, and data provenance treated as core experimental dimensions. The architecture integrates virtualized and physical edge nodes, digital twin coordination, time-windowed control, and combined stream processing with cloud-side aggregation analytics, enabling multi-domain CPS applications to share programmable infrastructure and flexibly deploy and compare control and monitoring strategies. Validation through 40 systematic experiments across geographically distributed deployments—spanning renewable energy community management and AirWatch monitoring use cases—demonstrates the framework’s effectiveness and generality in hybrid physical-virtual settings.

Cloud ContinuumCyber-Physical SystemsDistributed Experimentation

This study addresses the problem of edge service deployment failures in digital healthcare caused by network outages affecting centralized registries. To mitigate this, we propose a three-tier distributed registry architecture integrating remote public, MEC-private, and LAN-local tiers. By combining edge computing with Docker containerization, the proposed approach enables proximity-aware service orchestration and high-availability distribution. Experimental evaluations demonstrate that under network disconnection scenarios, private and local registries significantly outperform their public counterparts in terms of deployment latency, system load, and energy consumption. These findings indicate that the proposed architecture effectively enhances the fault tolerance and overall resilience of edge services in healthcare environments.

Digital HealthcareEdge-Cloud ContinuumNetwork Disruption

This study addresses the problem of Kubernetes infrastructure drift, where runtime states deviate from architectural intent, by proposing an editable living architecture model. This model explicitly maps runtime facts to architectural designs, supports bidirectional synchronization between textual and graphical views, and enables non-intrusive Architecture-as-Code management through periodic consistency checks. The proposed approach is implemented using a subset of the Archer KDL, a VS Code-based prototype tool, and snapshot restoration techniques. Experimental evaluations conducted on three representative applications validate the feasibility of the method, demonstrating strong performance in both snapshot restoration accuracy and inconsistency detection recall.

Architecture-as-CodeCloud-native architectureConformance checking

Hot Scholars

SG

Swaroop Ghosh

Pennsylvania State University
Emerging memory technologiesHardware securityQuantum computingQuantum machine learning
AH

Andrew Hornback

PhD Candidate, Computer Science
Artificial IntelligenceComputer ScienceDeep LearningDerivatives
JS

Jakub Szefer

Associate Professor of Electrical Engineering, Yale University
Computer SecurityHardware SecurityFPGA SecurityQuantum Computer Security