Score
Designs, builds, and integrates foundational systems and platforms that must reliably support very large scale in users, load, or data, including architecture, capacity planning, deployment automation, monitoring, resilience, and operational procedures. Analyzes performance, failure modes, scalability limits, and cost/maintenance trade‑offs to ensure the infrastructure meets availability, throughput, and evolvability requirements.
This study addresses the lack of systematic guidance for enterprise software teams in choosing between monolithic and microservices architectures. The work proposes a decision-making framework that integrates technical and organizational factors, evaluating the trade-offs of each architecture across dimensions such as scalability, reliability, deployment efficiency, and organizational complexity. The assessment is grounded in system scale, business requirements, operational maturity, and long-term maintainability. Through architectural pattern analysis, a structured evaluation model, and multiple case studies, the authors develop a practical selection methodology tailored to real-world engineering contexts. This approach offers enterprises clear architectural evolution pathways and actionable guidelines aligned with their developmental stages, thereby significantly enhancing the rationality and sustainability of system design decisions.
This work addresses three core challenges in hybrid computing involving foundation models (FMs) and symbolic programs: semantic misalignment, insufficient reliability, and scalability bottlenecks. Methodologically, we propose a capability-complementary hybrid computation offloading paradigm, underpinned by an infrastructure framework that supports dynamic task offloading and scheduling. The framework integrates task decomposition, resource-aware allocation, and adaptive optimization, while unifying formal verification of symbolic programs with FM-based inference interfaces. To our knowledge, this is the first approach to achieve organic synergy between FM-driven semantic understanding and the deterministic execution guarantees of symbolic programs. As a result, the system simultaneously achieves high accuracy, strong generalization, enhanced stability, and improved efficiency in large-scale data processing. This work establishes a foundational architectural blueprint for building efficient, reliable, and formally verifiable hybrid intelligent software systems.
This work addresses the challenge of resource allocation in geographically distributed and heterogeneous continuum computing infrastructures, where combinatorial explosion and limited generalization hinder effective deployment. To tackle this, the study introduces, for the first time, the pricing structures commonly found in Software-as-a-Service (SaaS) ecosystems into the resource allocation problem, formulating a unified, price-based representation of the configuration space. The authors propose PRIME, a pricing-aware analysis engine that efficiently searches for cost-optimal deployment configurations satisfying both functional and non-functional constraints. Leveraging synthetic infrastructure topologies and workload generation techniques, the project constructs a comprehensive dataset comprising 9,600 diverse scenarios, demonstrating that the proposed approach achieves both scalability and computational efficiency in complex, heterogeneous environments.
This study addresses the infrastructure complexity of cloud-edge-end协同 architectures, which has emerged as a major bottleneck hindering developer productivity and innovation. Through 101 semi-structured interviews across 86 organizations, this work empirically identifies deployment complexity and onboarding difficulty as core challenges. It proposes four architectural directions to mitigate these issues: Object-as-a-Service (unified object abstraction), internal developer platforms, declarative AI/ML pipelines, and lightweight edge runtimes. Findings indicate that high-level abstractions and automation significantly enhance developer experience—outweighing the impact of execution performance optimizations—and thereby establish a new paradigm for platform engineering and distributed system design.
In microservice-based cloud-native systems, auto-scaling effectiveness is fundamentally constrained by architectural design, implementation choices, and deployment practices across the software lifecycle—factors often overlooked by existing benchmarks, leading to misleading evaluations. This paper systematically classifies and identifies critical engineering challenges affecting scaling performance according to software lifecycle phases, introducing the “lifecycle-aware scaling design” paradigm—the first of its kind. Using the Sock-Shop benchmark, we comparatively evaluate five scaling strategies: threshold-based, control-theoretic, machine-learning-driven, black-box optimization, and dependency-aware approaches. Experimental results demonstrate that holistically integrating lifecycle considerations significantly improves scaling stability (37% reduction in metric volatility) and resource efficiency (22% higher CPU utilization), whereas neglecting them causes severe performance degradation. This work bridges the gap between algorithmic auto-scaling research and real-world engineering deployment, providing a foundational methodology for production-grade scaling.
This study addresses cascading failures caused by shared firmware in commercial multi-host network interface cards (NICs) and the operational challenges of hyperscale deployments by proposing fbnic, a system-level solution. Architecturally, it introduces physical isolation and driver-priority mechanisms, combined with sub-sled-granularity firmware upgrade orchestration and slice-level fault containment. Operationally, it establishes a hardware-in-the-loop (HIL) continuous integration pipeline alongside cross-layer fault attribution monitoring and an automated remediation toolchain. Deployed across hundreds of thousands of hosts, fbnic reduces unplanned unavailability by 12×, shortens mean time to repair by 37%, and decreases hardware replacement rates by 2.3×, demonstrating robust stability at scale.
Traditional high-availability clusters are often constrained by single points of failure and inefficient resource allocation, making it difficult to meet the continuous availability demands of enterprise-grade systems. This work proposes an integrated High-Availability Cluster (iHAC), which innovatively combines active-active and active-passive architectures to optimize load distribution and failover mechanisms. By harmonizing these approaches, iHAC enhances fault tolerance while significantly improving resource utilization. Simulation experiments conducted using Riverbed Modeler (OPNET) demonstrate that iHAC reduces the average HTTP page response time by over 40%—from 5 seconds to under 3 seconds—compared to conventional solutions. This improvement translates into markedly lower network latency and higher system throughput, underscoring the efficacy of the proposed architecture in real-world deployment scenarios.
This study addresses the lack of empirical evidence on how microservice topology influences system performance and energy efficiency. Leveraging the μBench framework, the authors construct six canonical topologies—including chain, mesh, hierarchical, fan-out, probabilistic, and parallel fan-out—and conduct standardized load experiments across service scales of 5, 10, and 20 instances. Comprehensive metrics such as throughput, response time, energy consumption, CPU utilization, and failure rate are systematically evaluated. The work presents the first multidimensional quantification of topology-specific energy-performance trade-offs, revealing convergence patterns under scaling: mesh exhibits the poorest efficiency, while hierarchical, chain, and fan-out topologies offer more balanced behavior. Notably, under CPU-intensive workloads, probabilistic and parallel fan-out topologies achieve superior energy efficiency at larger scales, providing empirical foundations for green microservice architecture design.
本文通过引入一个多维度度量框架,解决多尺度高性能计算中的资源管理和可持续性问题,指导现代工作负载的部署策略。
本文通过实证方法研究了CPU核心数、RAM量和磁盘带宽对Ceph分布式文件系统性能的影响,旨在寻找成本最优的基础设施规模。