Score
Designs, builds, and analyzes scalable, secure Microsoft Azure cloud infrastructures and platform solutions, covering compute, networking, identity, storage, data platform services, app and ML services, and service integration. Produces solution and enterprise architecture artifacts and patterns, including deployment and DevOps pipelines, integration designs, and operational governance for Azure-based applications and data workloads.
This study addresses the escalating complexity of software architectures driven by cloud-native paradigms, microservices, and AI integration by systematically reviewing literature from 2024 to 2025. Focusing on five key dimensions—architectural modeling, quality attributes, self-adaptation mechanisms, AI-assisted decision-making, and architectural evolution—the work employs a thematic synthesis approach to integrate, for the first time, AI-enabled architectural decisions with continuous governance perspectives. It emphasizes the synergy among multi-view modeling, domain-driven decomposition, and runtime observability. The review identifies critical research gaps, including the absence of standardized frameworks for trustworthy AI architectures, insufficient empirical validation, weak integration of security and privacy concerns, and limited investigation into edge and serverless contexts. These insights offer actionable pathways to enhance scalability, maintainability, and software sustainability.
This study addresses the practical disparities and co-evolution between high-performance computing (HPC) and edge computing architectures within the cloud continuum. It presents the first large-scale empirical analysis based on 396 real-world, production-grade AWS architectures. Methodologically, we propose a multidimensional, data-driven framework encompassing service topology identification, storage type classification, architectural complexity quantification, and ML service integration statistics. Results reveal systematic differences—and complementary patterns—between HPC and edge architectures across four dimensions: core service composition (e.g., EC2 versus Greengrass/Lambda), storage design paradigms (parallel file systems versus distributed lightweight caches), complexity distributions, and ML embedding strategies. This work delivers the first industry-scale architectural benchmark for the cloud continuum, providing empirically grounded insights and methodological foundations for cross-domain architecture design, resource optimization, and cloud-native convergence of HPC and edge computing.
This study addresses the infrastructure complexity of cloud-edge-end协同 architectures, which has emerged as a major bottleneck hindering developer productivity and innovation. Through 101 semi-structured interviews across 86 organizations, this work empirically identifies deployment complexity and onboarding difficulty as core challenges. It proposes four architectural directions to mitigate these issues: Object-as-a-Service (unified object abstraction), internal developer platforms, declarative AI/ML pipelines, and lightweight edge runtimes. Findings indicate that high-level abstractions and automation significantly enhance developer experience—outweighing the impact of execution performance optimizations—and thereby establish a new paradigm for platform engineering and distributed system design.
In cloud-native systems, Kubernetes’ declarative configurations (YAML/Helm) impede architectural understanding, hindering developer and operator productivity. This paper introduces the first user-study-driven visualization framework that automatically and semantically faithfully maps raw Kubernetes resource specifications to interpretable architecture diagrams. Methodologically, it integrates a custom domain-specific language (DSL), the Kubernetes client API, and a graph-based encoding strategy to enable zero-intrusion integration and low-friction embedding into DevOps pipelines. Evaluated on three real-world systems, the tool accelerates architectural comprehension by 42% on average and improves modeling accuracy—reducing modeling errors by 68%. It has been adopted by the CNCF ecosystem as a recommended visualization tool. The core contributions are: (1) the first visualization generation paradigm explicitly designed for Kubernetes semantics; and (2) a solution that jointly ensures precision, scalability, and engineering practicality.
This study addresses the lack of systematic optimization in cloud data pipelines with respect to cost, execution time, and resource utilization, particularly in multi-tenant and industrial settings where research remains limited. Through a comprehensive systematic literature review, the work establishes a unified classification framework for optimization objectives that encompasses both single- and multi-cloud environments as well as batch and stream processing paradigms. The analysis synthesizes existing approaches and identifies critical research gaps, including insufficient support for multi-tenancy, inadequate multi-cloud coordination, and a scarcity of real-world deployment validation. By clarifying the core objectives and technical pathways for optimizing cloud data pipelines, this paper provides a theoretical foundation and clear direction for future research in this domain.
This study addresses prolonged task completion times, low resource utilization, and high resource release latency in Docker/Kubernetes containers on cloud-native platforms running compute-intensive workloads (e.g., big data and deep learning). We systematically evaluate the performance impact of diverse resource scheduling strategies through system-level monitoring—leveraging cgroups and metrics-server—and multi-workload stress testing. For the first time, we empirically quantify how key resource configurations significantly affect task completion time (±79.4% variation) and resource release latency (+116.7% degradation). Based on these findings, we propose an evidence-driven configuration optimization paradigm that reduces maximum task completion time by up to 79.4% and precisely identifies configuration bottlenecks responsible for latency. Our results provide reproducible, transferable empirical foundations for resource management tuning and deployment decisions in cloud-native environments.
This work addresses the challenge of detecting multi-stage attacks that traverse trust boundaries in cloud deployments—threats often missed by conventional security tools due to their inability to model holistic system architecture and runtime behavioral deviations. The authors propose a novel approach that integrates static configuration analysis with runtime network flow observation to automatically construct a platform-agnostic architectural abstraction reflecting the system’s true state, including components, domains, interfaces, policies, and data flows. Building upon this representation, the method enables continuous, architecture-level threat modeling. It is the first to support automated architecture inference and threat detection across bare-metal, Kubernetes, and cloud environments. Evaluated on supply chain systems incorporating machine learning (ML) components, the approach successfully identified all 17 classes of injection threats—including ML-specific threats—substantially outperforming existing tools, which cover only 6–47% of these threats and fail entirely to detect ML-related ones.
Enterprise cloud environments are frequently exposed to security threats due to misconfigurations, excessive permissions, and fragmented security tooling, compounded by the absence of unified, coordinated protection across Kubernetes, OpenStack, and Infrastructure-as-Code (IaC) platforms. This work proposes the first open-source microservices-based security framework that uniquely integrates identity governance, multi-platform configuration auditing, runtime threat detection, and automated IaC remediation into a single closed-loop system. Designed with standardized REST/gRPC interfaces and scalable for medium-to-large deployments, the framework synergistically combines Falco, ELK, Terraform, Checkov, and OPA. In enterprise evaluations, it reduced vulnerability assessment time from 120 to 18 minutes, achieved a false positive rate below 5%, decreased security incidents by 62%, and lowered operational costs by approximately 40%, all while being released under the Apache 2.0 license.
This work addresses the challenges of resource utilization and operational efficiency in microservice architectures by proposing a performance-metric-driven automated framework that intelligently determines the optimal deployment strategy for individual microservices between Infrastructure-as-a-Service (IaaS) and Function-as-a-Service (FaaS). By analyzing intrinsic microservice characteristics, the framework enables a scalable and reproducible migration from conventional IaaS deployments to a hybrid IaaS+FaaS model. Experimental evaluation on two real-world applications demonstrates that the approach accurately identifies microservices well-suited for serverless execution, significantly improving both deployment efficiency and resource utilization. Furthermore, the study clarifies the respective applicability boundaries and advantages of different cloud service models, offering practical guidance for architecture design in heterogeneous cloud environments.
This study addresses the coordination challenges arising from independent control across cloud, high-performance computing (HPC), and edge AI infrastructures. Conceptualizing the AI platform as a "system of systems," this work proposes an architectural paradigm characterized by usage fusion and federated control. Methodologically, it adopts a systems engineering framework that achieves cross-domain coordination through interface contracts while preserving native control planes. The approach incorporates boundary testing, responsibility models, and seven integration facets, leveraging interface mapping, policy contexts, and operational evidence to guide integration design. The primary contribution lies in establishing a unified framework for evaluating interoperability, governance capabilities, and fault isolation, thereby delineating clear directions for future research.
This study addresses the lack of systematic guidance for enterprise software teams in choosing between monolithic and microservices architectures. The work proposes a decision-making framework that integrates technical and organizational factors, evaluating the trade-offs of each architecture across dimensions such as scalability, reliability, deployment efficiency, and organizational complexity. The assessment is grounded in system scale, business requirements, operational maturity, and long-term maintainability. Through architectural pattern analysis, a structured evaluation model, and multiple case studies, the authors develop a practical selection methodology tailored to real-world engineering contexts. This approach offers enterprises clear architectural evolution pathways and actionable guidelines aligned with their developmental stages, thereby significantly enhancing the rationality and sustainability of system design decisions.