Score
Designs, builds, and analyzes the foundational hardware, software, and operational systems that host, connect, and run applications and services — including provisioning, configuration, deployment, orchestration, networking, storage, monitoring, and fault-tolerance mechanisms. Implements automation, capacity planning, performance optimization, security controls, and observability to ensure scalable, reliable, and maintainable system operation.
This study addresses cascading failures caused by shared firmware in commercial multi-host network interface cards (NICs) and the operational challenges of hyperscale deployments by proposing fbnic, a system-level solution. Architecturally, it introduces physical isolation and driver-priority mechanisms, combined with sub-sled-granularity firmware upgrade orchestration and slice-level fault containment. Operationally, it establishes a hardware-in-the-loop (HIL) continuous integration pipeline alongside cross-layer fault attribution monitoring and an automated remediation toolchain. Deployed across hundreds of thousands of hosts, fbnic reduces unplanned unavailability by 12×, shortens mean time to repair by 37%, and decreases hardware replacement rates by 2.3×, demonstrating robust stability at scale.
Cloud computing systems exhibit high design and operational complexity, heavily reliant on manual intervention. To address this, we propose an intent-driven automation paradigm that unifies functional requirements and runtime constraints into a single abstraction, enabling autonomous system management across the entire lifecycle—design, implementation, operation, and evolution. We formally define “intent” as the core abstraction for cloud systems and introduce a four-component framework comprising intent modeling, parsing, verification, and execution. This framework integrates formal methods, domain-specific languages (DSLs), constraint solving, and feedback-based closed-loop control. Our work establishes foundational theory for autonomous systems, delivers a scalable intent-driven roadmap, and provides a methodological foundation and community-aligned framework for cloud-native and AI-native infrastructure. (138 words)
This study addresses the coordination challenges arising from independent control across cloud, high-performance computing (HPC), and edge AI infrastructures. Conceptualizing the AI platform as a "system of systems," this work proposes an architectural paradigm characterized by usage fusion and federated control. Methodologically, it adopts a systems engineering framework that achieves cross-domain coordination through interface contracts while preserving native control planes. The approach incorporates boundary testing, responsibility models, and seven integration facets, leveraging interface mapping, policy contexts, and operational evidence to guide integration design. The primary contribution lies in establishing a unified framework for evaluating interoperability, governance capabilities, and fault isolation, thereby delineating clear directions for future research.
This study addresses the lack of systematic guidance for enterprise software teams in choosing between monolithic and microservices architectures. The work proposes a decision-making framework that integrates technical and organizational factors, evaluating the trade-offs of each architecture across dimensions such as scalability, reliability, deployment efficiency, and organizational complexity. The assessment is grounded in system scale, business requirements, operational maturity, and long-term maintainability. Through architectural pattern analysis, a structured evaluation model, and multiple case studies, the authors develop a practical selection methodology tailored to real-world engineering contexts. This approach offers enterprises clear architectural evolution pathways and actionable guidelines aligned with their developmental stages, thereby significantly enhancing the rationality and sustainability of system design decisions.
To address insufficient resilience of complex systems under heterogeneous hardware environments, this paper proposes a fault-adaptive software deployment and redundancy configuration optimization method. We construct a system-level resilience state-space model and introduce a novel equivalence relation to enable quotient-space-based state-space reduction, significantly compressing the state space. Subsequently, we integrate formal model checking with strategy synthesis to automatically derive both an initial deployment configuration and dynamic reconfiguration policies that satisfy multi-level resilience requirements. Our key contributions are: (i) a new equivalence relation enabling efficient, semantics-preserving state-space reduction; and (ii) end-to-end automated synthesis of fault-response and recovery strategies. Experimental evaluation on an autonomous driving system model demonstrates that our approach substantially improves fault recovery latency and system availability, while supporting real-time resilience assurance.
本文解决了RISC-V系统中FPGA生命周期管理缺乏通用模型的问题,通过利用标准Linux功能提出了一种与主机无关的控制平面架构。
This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.
Serverless computing has matured into an effective execution model for edge cloud environments, enabling function level decomposition, demand driven scaling, and workflow execution across stable, well provisioned infrastructure. This success motivates extending it to the edge cloud space continuum, where Low Earth Orbit (LEO) constellations are increasingly explored as distributed compute substrates. However, existing serverless orchestration is not directly applicable in this setting, where LEO systems impose time varying contact graphs, intermittent link availability, and strict feasibility constraints on energy, memory, communication, and operational cost. This article identifies ten broken assumptions in existing serverless orchestration and organizes them into three core challenges: spatiotemporal execution over dynamic graphs, constraint aware function placement and scaling, and correctness and progress under decentralized and delayed state. It then proposes an architecture that enables robust and efficient serverless execution across the continuum, grounded in these challenges and demonstrated through a representative flood response use case.
This work addresses the persistent challenge of inconsistent development and execution environments faced by researchers operating across heterogeneous computing platforms—ranging from laptops and workstations to supercomputers and cloud infrastructures. To overcome this, the authors propose a modular and portable software ecosystem featuring a unified command-line interface that enables seamless orchestration and execution of scientific workflows. The system ensures cross-platform consistency, reproducibility, and scalability, thereby streamlining computational research across diverse hardware configurations. Its practical efficacy has been demonstrated through successful integration into the plan4res project under the European Union’s Horizon 2020 initiative, where it effectively supported complex, large-scale scientific workflows in varied computing environments.
This study addresses the infrastructure complexity of cloud-edge-end协同 architectures, which has emerged as a major bottleneck hindering developer productivity and innovation. Through 101 semi-structured interviews across 86 organizations, this work empirically identifies deployment complexity and onboarding difficulty as core challenges. It proposes four architectural directions to mitigate these issues: Object-as-a-Service (unified object abstraction), internal developer platforms, declarative AI/ML pipelines, and lightweight edge runtimes. Findings indicate that high-level abstractions and automation significantly enhance developer experience—outweighing the impact of execution performance optimizations—and thereby establish a new paradigm for platform engineering and distributed system design.