Score
Designs, specifies, and analyzes system and solution architectures that scale in performance, capacity, and availability, covering distributed systems, cloud and platform architectures, microservices, and hardware–software co-designed systems. Creates reference and solution architectures, integration strategies, deployment and fault‑tolerance patterns, and scalability plans that define components, interfaces, trade‑offs, and operational requirements to meet throughput, latency, reliability, and capacity goals.
This study addresses the escalating complexity of software architectures driven by cloud-native paradigms, microservices, and AI integration by systematically reviewing literature from 2024 to 2025. Focusing on five key dimensions—architectural modeling, quality attributes, self-adaptation mechanisms, AI-assisted decision-making, and architectural evolution—the work employs a thematic synthesis approach to integrate, for the first time, AI-enabled architectural decisions with continuous governance perspectives. It emphasizes the synergy among multi-view modeling, domain-driven decomposition, and runtime observability. The review identifies critical research gaps, including the absence of standardized frameworks for trustworthy AI architectures, insufficient empirical validation, weak integration of security and privacy concerns, and limited investigation into edge and serverless contexts. These insights offer actionable pathways to enhance scalability, maintainability, and software sustainability.
This study addresses the lack of systematic guidance for enterprise software teams in choosing between monolithic and microservices architectures. The work proposes a decision-making framework that integrates technical and organizational factors, evaluating the trade-offs of each architecture across dimensions such as scalability, reliability, deployment efficiency, and organizational complexity. The assessment is grounded in system scale, business requirements, operational maturity, and long-term maintainability. Through architectural pattern analysis, a structured evaluation model, and multiple case studies, the authors develop a practical selection methodology tailored to real-world engineering contexts. This approach offers enterprises clear architectural evolution pathways and actionable guidelines aligned with their developmental stages, thereby significantly enhancing the rationality and sustainability of system design decisions.
This study addresses the practical disparities and co-evolution between high-performance computing (HPC) and edge computing architectures within the cloud continuum. It presents the first large-scale empirical analysis based on 396 real-world, production-grade AWS architectures. Methodologically, we propose a multidimensional, data-driven framework encompassing service topology identification, storage type classification, architectural complexity quantification, and ML service integration statistics. Results reveal systematic differences—and complementary patterns—between HPC and edge architectures across four dimensions: core service composition (e.g., EC2 versus Greengrass/Lambda), storage design paradigms (parallel file systems versus distributed lightweight caches), complexity distributions, and ML embedding strategies. This work delivers the first industry-scale architectural benchmark for the cloud continuum, providing empirically grounded insights and methodological foundations for cross-domain architecture design, resource optimization, and cloud-native convergence of HPC and edge computing.
In microservice-based cloud-native systems, auto-scaling effectiveness is fundamentally constrained by architectural design, implementation choices, and deployment practices across the software lifecycle—factors often overlooked by existing benchmarks, leading to misleading evaluations. This paper systematically classifies and identifies critical engineering challenges affecting scaling performance according to software lifecycle phases, introducing the “lifecycle-aware scaling design” paradigm—the first of its kind. Using the Sock-Shop benchmark, we comparatively evaluate five scaling strategies: threshold-based, control-theoretic, machine-learning-driven, black-box optimization, and dependency-aware approaches. Experimental results demonstrate that holistically integrating lifecycle considerations significantly improves scaling stability (37% reduction in metric volatility) and resource efficiency (22% higher CPU utilization), whereas neglecting them causes severe performance degradation. This work bridges the gap between algorithmic auto-scaling research and real-world engineering deployment, providing a foundational methodology for production-grade scaling.
This study addresses the lack of empirical evidence on how microservice topology influences system performance and energy efficiency. Leveraging the μBench framework, the authors construct six canonical topologies—including chain, mesh, hierarchical, fan-out, probabilistic, and parallel fan-out—and conduct standardized load experiments across service scales of 5, 10, and 20 instances. Comprehensive metrics such as throughput, response time, energy consumption, CPU utilization, and failure rate are systematically evaluated. The work presents the first multidimensional quantification of topology-specific energy-performance trade-offs, revealing convergence patterns under scaling: mesh exhibits the poorest efficiency, while hierarchical, chain, and fan-out topologies offer more balanced behavior. Notably, under CPU-intensive workloads, probabilistic and parallel fan-out topologies achieve superior energy efficiency at larger scales, providing empirical foundations for green microservice architecture design.
Microservice system developers lack empirical evidence regarding the types, root causes, and remediation strategies of recurring issues. Method: We adopt a mixed-methods approach—quantitatively analyzing 2,641 open-source issues, qualitatively interviewing 15 practitioners, and conducting a global survey with 150 practitioners. Contribution/Results: We introduce the first comprehensive, domain-specific three-level taxonomy (“Issue–Cause–Solution”) for microservices. We identify five high-frequency issue domains—including technical debt, CI/CD pipeline failures, and exception handling—and three predominant root causes, notably generic programming errors. From our analysis, we distill 177 actionable, context-aware remediation strategies. This work establishes an empirical foundation for microservice fault diagnosis and mitigation, delivers practical guidance for industry practitioners, and pinpoints critical research directions for next-generation microservice engineering.
This study addresses the longstanding fragmentation in microservice energy efficiency research, which has been siloed across runtime, infrastructure, and architectural layers, lacking a unified lifecycle perspective and consistent measurement methodology. Employing Kitchenham’s systematic literature review approach—augmented by searches across four major databases and snowballing techniques—the authors analyze 40 core studies to integrate multidimensional viewpoints for the first time. Their synthesis reveals an overwhelming emphasis on runtime optimizations, such as scheduling and resource management, typically relying on coarse-grained monitoring and model-based estimations, while largely neglecting energy-aware integration during architectural design and fine-grained measurement practices. The work establishes energy efficiency as a critical cross-cutting architectural attribute throughout the microservice lifecycle and underscores the urgent need for early-design support and a unified measurement framework.
This study addresses the infrastructure complexity of cloud-edge-end协同 architectures, which has emerged as a major bottleneck hindering developer productivity and innovation. Through 101 semi-structured interviews across 86 organizations, this work empirically identifies deployment complexity and onboarding difficulty as core challenges. It proposes four architectural directions to mitigate these issues: Object-as-a-Service (unified object abstraction), internal developer platforms, declarative AI/ML pipelines, and lightweight edge runtimes. Findings indicate that high-level abstractions and automation significantly enhance developer experience—outweighing the impact of execution performance optimizations—and thereby establish a new paradigm for platform engineering and distributed system design.
This study addresses the coordination challenges arising from independent control across cloud, high-performance computing (HPC), and edge AI infrastructures. Conceptualizing the AI platform as a "system of systems," this work proposes an architectural paradigm characterized by usage fusion and federated control. Methodologically, it adopts a systems engineering framework that achieves cross-domain coordination through interface contracts while preserving native control planes. The approach incorporates boundary testing, responsibility models, and seven integration facets, leveraging interface mapping, policy contexts, and operational evidence to guide integration design. The primary contribution lies in establishing a unified framework for evaluating interoperability, governance capabilities, and fault isolation, thereby delineating clear directions for future research.
This work addresses the challenges of resource utilization and operational efficiency in microservice architectures by proposing a performance-metric-driven automated framework that intelligently determines the optimal deployment strategy for individual microservices between Infrastructure-as-a-Service (IaaS) and Function-as-a-Service (FaaS). By analyzing intrinsic microservice characteristics, the framework enables a scalable and reproducible migration from conventional IaaS deployments to a hybrid IaaS+FaaS model. Experimental evaluation on two real-world applications demonstrates that the approach accurately identifies microservices well-suited for serverless execution, significantly improving both deployment efficiency and resource utilization. Furthermore, the study clarifies the respective applicability boundaries and advantages of different cloud service models, offering practical guidance for architecture design in heterogeneous cloud environments.
This work addresses the persistent challenge of inconsistent development and execution environments faced by researchers operating across heterogeneous computing platforms—ranging from laptops and workstations to supercomputers and cloud infrastructures. To overcome this, the authors propose a modular and portable software ecosystem featuring a unified command-line interface that enables seamless orchestration and execution of scientific workflows. The system ensures cross-platform consistency, reproducibility, and scalability, thereby streamlining computational research across diverse hardware configurations. Its practical efficacy has been demonstrated through successful integration into the plan4res project under the European Union’s Horizon 2020 initiative, where it effectively supported complex, large-scale scientific workflows in varied computing environments.