design scalable systems

Designs, specifies, and analyzes system and solution architectures that scale in performance, capacity, and availability, covering distributed systems, cloud and platform architectures, microservices, and hardware–software co-designed systems. Creates reference and solution architectures, integration strategies, deployment and fault‑tolerance patterns, and scalability plans that define components, interfaces, trade‑offs, and operational requirements to meet throughput, latency, reliability, and capacity goals.

designscalablesystems

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.14
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$202K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the lack of systematic guidance for enterprise software teams in choosing between monolithic and microservices architectures. The work proposes a decision-making framework that integrates technical and organizational factors, evaluating the trade-offs of each architecture across dimensions such as scalability, reliability, deployment efficiency, and organizational complexity. The assessment is grounded in system scale, business requirements, operational maturity, and long-term maintainability. Through architectural pattern analysis, a structured evaluation model, and multiple case studies, the authors develop a practical selection methodology tailored to real-world engineering contexts. This approach offers enterprises clear architectural evolution pathways and actionable guidelines aligned with their developmental stages, thereby significantly enhancing the rationality and sustainability of system design decisions.

MicroservicesMonolithic ArchitectureOrganizational Complexity

An Analysis of HPC and Edge Architectures in the Cloud

Aug 02, 2025
SS
Steven Santillan
🏛️ Escuela Superior Politécnica del Litoral | ESPOL

This study addresses the practical disparities and co-evolution between high-performance computing (HPC) and edge computing architectures within the cloud continuum. It presents the first large-scale empirical analysis based on 396 real-world, production-grade AWS architectures. Methodologically, we propose a multidimensional, data-driven framework encompassing service topology identification, storage type classification, architectural complexity quantification, and ML service integration statistics. Results reveal systematic differences—and complementary patterns—between HPC and edge architectures across four dimensions: core service composition (e.g., EC2 versus Greengrass/Lambda), storage design paradigms (parallel file systems versus distributed lightweight caches), complexity distributions, and ML embedding strategies. This work delivers the first industry-scale architectural benchmark for the cloud continuum, providing empirically grounded insights and methodological foundations for cross-domain architecture design, resource optimization, and cloud-native convergence of HPC and edge computing.

Analyze HPC and edge architectures in AWS cloud deploymentsAssess architectural complexity and machine learning services usageInvestigate AWS services prevalence and storage systems used

Key Considerations for Auto-Scaling: Lessons from Benchmark Microservices

Oct 02, 2025
MD
Majid Dashtbani
🏛️ University of Waterloo

In microservice-based cloud-native systems, auto-scaling effectiveness is fundamentally constrained by architectural design, implementation choices, and deployment practices across the software lifecycle—factors often overlooked by existing benchmarks, leading to misleading evaluations. This paper systematically classifies and identifies critical engineering challenges affecting scaling performance according to software lifecycle phases, introducing the “lifecycle-aware scaling design” paradigm—the first of its kind. Using the Sock-Shop benchmark, we comparatively evaluate five scaling strategies: threshold-based, control-theoretic, machine-learning-driven, black-box optimization, and dependency-aware approaches. Experimental results demonstrate that holistically integrating lifecycle considerations significantly improves scaling stability (37% reduction in metric volatility) and resource efficiency (22% higher CPU utilization), whereas neglecting them causes severe performance degradation. This work bridges the gap between algorithmic auto-scaling research and real-world engineering deployment, providing a foundational methodology for production-grade scaling.

Addressing overlooked lifecycle considerations in microservice auto-scaling benchmarksClassifying auto-scaling challenges across Architecture, Implementation, and Deployment phasesValidating how lifecycle awareness improves autoscaler stability and efficiency

This study addresses the lack of empirical evidence on how microservice topology influences system performance and energy efficiency. Leveraging the μBench framework, the authors construct six canonical topologies—including chain, mesh, hierarchical, fan-out, probabilistic, and parallel fan-out—and conduct standardized load experiments across service scales of 5, 10, and 20 instances. Comprehensive metrics such as throughput, response time, energy consumption, CPU utilization, and failure rate are systematically evaluated. The work presents the first multidimensional quantification of topology-specific energy-performance trade-offs, revealing convergence patterns under scaling: mesh exhibits the poorest efficiency, while hierarchical, chain, and fan-out topologies offer more balanced behavior. Notably, under CPU-intensive workloads, probabilistic and parallel fan-out topologies achieve superior energy efficiency at larger scales, providing empirical foundations for green microservice architecture design.

architectural topologycloud-nativeenergy efficiency

Understanding the Issues, Their Causes and Solutions in Microservices Systems: An Empirical Study

Feb 03, 2023
MW
Muhammad Waseem
🏛️ Wuhan University | Lancaster University Leipzig | University of Oulu | RMIT University | Tampere University | University of Jyväskylä

Microservice system developers lack empirical evidence regarding the types, root causes, and remediation strategies of recurring issues. Method: We adopt a mixed-methods approach—quantitatively analyzing 2,641 open-source issues, qualitatively interviewing 15 practitioners, and conducting a global survey with 150 practitioners. Contribution/Results: We introduce the first comprehensive, domain-specific three-level taxonomy (“Issue–Cause–Solution”) for microservices. We identify five high-frequency issue domains—including technical debt, CI/CD pipeline failures, and exception handling—and three predominant root causes, notably generic programming errors. From our analysis, we distill 177 actionable, context-aware remediation strategies. This work establishes an empirical foundation for microservice fault diagnosis and mitigation, delivers practical guidance for industry practitioners, and pinpoints critical research directions for next-generation microservice engineering.

Analyzing root causes behind microservices failuresDeveloping comprehensive solutions for microservices problemsIdentifying dominant issues in microservices systems

Latest Papers

What's happening recently
View more

This study addresses the longstanding fragmentation in microservice energy efficiency research, which has been siloed across runtime, infrastructure, and architectural layers, lacking a unified lifecycle perspective and consistent measurement methodology. Employing Kitchenham’s systematic literature review approach—augmented by searches across four major databases and snowballing techniques—the authors analyze 40 core studies to integrate multidimensional viewpoints for the first time. Their synthesis reveals an overwhelming emphasis on runtime optimizations, such as scheduling and resource management, typically relying on coarse-grained monitoring and model-based estimations, while largely neglecting energy-aware integration during architectural design and fine-grained measurement practices. The work establishes energy efficiency as a critical cross-cutting architectural attribute throughout the microservice lifecycle and underscores the urgent need for early-design support and a unified measurement framework.

architectural concernenergy efficiencymicroservice architectures

This study addresses the infrastructure complexity of cloud-edge-end协同 architectures, which has emerged as a major bottleneck hindering developer productivity and innovation. Through 101 semi-structured interviews across 86 organizations, this work empirically identifies deployment complexity and onboarding difficulty as core challenges. It proposes four architectural directions to mitigate these issues: Object-as-a-Service (unified object abstraction), internal developer platforms, declarative AI/ML pipelines, and lightweight edge runtimes. Findings indicate that high-level abstractions and automation significantly enhance developer experience—outweighing the impact of execution performance optimizations—and thereby establish a new paradigm for platform engineering and distributed system design.

cloud-edge infrastructuredeveloper productivitydistributed computing

This study addresses the coordination challenges arising from independent control across cloud, high-performance computing (HPC), and edge AI infrastructures. Conceptualizing the AI platform as a "system of systems," this work proposes an architectural paradigm characterized by usage fusion and federated control. Methodologically, it adopts a systems engineering framework that achieves cross-domain coordination through interface contracts while preserving native control planes. The approach incorporates boundary testing, responsibility models, and seven integration facets, leveraging interface mapping, policy contexts, and operational evidence to guide integration design. The primary contribution lies in establishing a unified framework for evaluating interoperability, governance capabilities, and fault isolation, thereby delineating clear directions for future research.

Cloud-HPC-Edge AI PlatformsComposable IntegrationFederated Control

This work addresses the challenges of resource utilization and operational efficiency in microservice architectures by proposing a performance-metric-driven automated framework that intelligently determines the optimal deployment strategy for individual microservices between Infrastructure-as-a-Service (IaaS) and Function-as-a-Service (FaaS). By analyzing intrinsic microservice characteristics, the framework enables a scalable and reproducible migration from conventional IaaS deployments to a hybrid IaaS+FaaS model. Experimental evaluation on two real-world applications demonstrates that the approach accurately identifies microservices well-suited for serverless execution, significantly improving both deployment efficiency and resource utilization. Furthermore, the study clarifies the respective applicability boundaries and advantages of different cloud service models, offering practical guidance for architecture design in heterogeneous cloud environments.

Cloud DeploymentFaaSIaaS

This work addresses the persistent challenge of inconsistent development and execution environments faced by researchers operating across heterogeneous computing platforms—ranging from laptops and workstations to supercomputers and cloud infrastructures. To overcome this, the authors propose a modular and portable software ecosystem featuring a unified command-line interface that enables seamless orchestration and execution of scientific workflows. The system ensures cross-platform consistency, reproducibility, and scalability, thereby streamlining computational research across diverse hardware configurations. Its practical efficacy has been demonstrated through successful integration into the plan4res project under the European Union’s Horizon 2020 initiative, where it effectively supported complex, large-scale scientific workflows in varied computing environments.

computational workflowsportablereproducible

Hot Scholars

MG

Minyi Guo

IEEE Fellow, Chair Professor, Shanghai Jiao Tong University
Parallel ComputingCompiler OptimizationCloud ComputingNetworking
YD

Yilun Du

Harvard University
Artificial IntelligenceMachine LearningRoboticsComputer Vision
XW

Xinggang Wang

Professor, Huazhong University of Science and Technology
Artificial IntelligenceComputer VisionAutonomous DrivingObject Detection
FS

Frédéric Suter

Oak Ridge National Laboratory, IEEE Senior member
Computer ScienceWorkflowSchedulingSimulation
FW

Feiyi Wang

Distinguished Research Scientist & Group Leader, Analytics and AI Methods at Scale, NCCS/ORNL
HPCAI for Science at Scale