Score
Designs, builds, and analyzes the architectures, operational practices, tooling, and test frameworks that ensure systems, services, platforms, and data pipelines meet specified availability, durability, latency, and fault‑tolerance objectives under normal and adverse conditions. This work includes defining reliability requirements and SLOs, performing reliability modeling and analysis, conducting resilience and failure‑injection testing, implementing monitoring, alerting and runbooks, and improving production reliability through incident response, postmortems, capacity planning, and automated mitigation.
To address the challenges of standardizing Site Reliability Engineering (SRE) practices in heterogeneous environments and balancing system reliability with development agility, this paper proposes a customizable SRE process framework. The framework integrates automated operations, multidimensional observability (metrics, logs, traces), error-budget-driven governance, standardized incident response, and progressive delivery (canary and blue-green deployments). It is designed for cross-technology-stack adaptability, enabling contextual implementation of core SRE principles. Evaluated in production systems, the framework reduced mean time to recovery by 42%, decreased unplanned outages by 67%, lowered operational staffing requirements by 35%, and achieved 99.99% service availability. Its primary contribution is the first methodology for customizing SRE processes specifically for heterogeneous environments, empirically demonstrating synergistic improvements in both system reliability and operational efficiency.
This work addresses the unreliability of developer productivity dashboards, which often stems from ad hoc scripts that introduce undetected silent data gaps, eroding organizational trust. To resolve this, we propose a robust ELT pipeline grounded in DAG-based orchestration and the Medallion architecture, decoupling data extraction from transformation to preserve the immutability of raw data. Our approach introduces a state-driven dependency scheduling mechanism and, for the first time, treats metric pipelines as production-grade distributed systems. We emphasize the critical role of immutable raw history in enabling reliable metric redefinition. This methodology significantly enhances data reliability and freshness while effectively eliminating silent failures, thereby restoring organizational confidence in DevOps metrics.
本文提出了一种服务健康工程方法,通过结合遥测、工作流完成情况等手段来检测分布式系统中的静默故障和异步工作停滞问题。
In cloud-native systems, alert rules frequently suffer from false positives and false negatives due to the absence of design-phase validation, while existing tools lack systematic support for alert testing. To address this, we propose the “Alert-as-Experiment” paradigm—the first adaptation of the observability experimentation framework OXN to early-stage alert rule validation. Our approach enables closed-loop, development-time testing and continuous calibration of alert logic via simulated execution, synthetic observation data injection, and real-world scenario replay. It supports parameter tuning and repeatable verification of alert-triggering behavior, shifting alert engineering from empirical practice toward a testable, verifiable, and systematic discipline. Empirical evaluation demonstrates significant reductions in both false positive and false negative rates in production environments, alongside improved fault response latency and system maintainability.
This paper addresses the insufficient integration of early-system safety analysis with Model-Based Systems Engineering (MBSE). It comparatively evaluates three functional safety analysis methods—Failure Mode and Effects Analysis (FMEA), Functional Hazard Assessment (FHA), and Fault Feedback and Impact Propagation (FFIP)—and identifies FFIP as superior for detecting emergent behaviors, second-order effects, and fault propagation. Subsequently, it systematically reviews existing MBSE integration practices, categorizing them into four approaches: model transformation, custom algorithm development, built-in toolkits, and manual modeling. The study reveals that current integration efforts are predominantly focused on FMEA, while FHA and FFIP remain in exploratory stages, hindered by the absence of a unified framework and standardized guidelines. To bridge this gap, the paper proposes a novel, full-lifecycle safety analysis integration paradigm aligned with digital engineering transformation—enabling traceable, executable, and evolvable model-driven safety verification.
论文提出一种混合关键性架构框架,通过硬件隔离安全监控和健康向量等方法确保无人机群在关键任务中的安全性和可靠性。
Automotive electronic control units (ECUs) are intricate systems with hundreds of individual functions, numerous software components, and multiple interdependent tasks. A prevalent structural pattern in these systems are so-called cause-effect chains. While significant research efforts have been dedicated to the temporal analysis and optimization of these chains, particularly minimizing data age and function response times, other crucial non-functional properties remain relatively underexplored. In particular, the safety integrity level (SIL) classification substantially influences the system design by determining task colocation strategies. Improper sharing of functions or interweaving tasks with different safety levels can compromise the integrity of critical functions. Additionally, AUTOSAR basic software (BSW) (e.g. OS, runtime environment, communication stacks, or diagnostics) introduces complexity that varies based on task characteristics and SIL categories. Furthermore, memory requirements present another critical challenge, given the diversity of memory architectures and SIL-specific dependencies that strongly constrain task allocations. This paper thoroughly characterizes a real-world automotive application, describing an automotive application based on SIL constraints, the impact of basic software, and memory requirements. In this context, the Driverator configuration framework is introduced for scalable system analysis.
Traditional high-availability clusters are often constrained by single points of failure and inefficient resource allocation, making it difficult to meet the continuous availability demands of enterprise-grade systems. This work proposes an integrated High-Availability Cluster (iHAC), which innovatively combines active-active and active-passive architectures to optimize load distribution and failover mechanisms. By harmonizing these approaches, iHAC enhances fault tolerance while significantly improving resource utilization. Simulation experiments conducted using Riverbed Modeler (OPNET) demonstrate that iHAC reduces the average HTTP page response time by over 40%—from 5 seconds to under 3 seconds—compared to conventional solutions. This improvement translates into markedly lower network latency and higher system throughput, underscoring the efficacy of the proposed architecture in real-world deployment scenarios.
本文介绍了Meta为解决大规模系统持续部署中速度与可靠性之间的矛盾,通过构建服务健康检查器进行自动回滚等方法保障部署安全。
This study addresses the growing challenges traditional failure analysis methods face in the era of advanced packaging technologies—such as chiplets, hybrid bonding, and 3D stacking—by conducting an anonymous global survey of over 100 semiconductor design, packaging, and failure analysis organizations. The findings reveal that 69% of respondents prioritize heterogeneous integration products (mean importance score: 7.92/10), while 54% identify hybrid bonding as the most analytically challenging technique. A strong consensus emerges around the need for standardized data formats, with 83% of participants advocating for unified protocols, and high-resolution non-destructive imaging garners substantial support (mean score: 8.18/10). The research systematically identifies critical pain points in sample preparation and 3D structural inspection, offering empirical insights to guide industry standardization and technological innovation.