Score
Designs, implements, and interprets evaluation protocols, statistical models, and tests that measure and estimate system or model reliability, calibration, and robustness (including adversarial vulnerability), and that quantify failure modes, outage/failure probabilities, and reproducibility across datasets. Builds reliability-aware metrics, calibration analyses, robustness assessments, and reliability models to assess trade-offs under failure conditions and to produce actionable reliability assessments for engineering and design decisions.
Contemporary AI systems are rapidly advancing toward transformative capabilities, necessitating safety evaluation methodologies that transcend conventional static benchmarks. Method: We propose a novel three-dimensional safety assessment framework—“Capability–Propensity–Control”—that systematically integrates measurement targets (e.g., deception capability, power-seeking propensity, adversarial robustness), measurement modalities (behavioral testing, internal analysis), and governance mapping. This framework overcomes limitations of static benchmarking by enabling dynamic, multi-layered evaluation. Contribution/Results: We introduce the first unified taxonomy covering the full stack of AI safety assessment; identify critical evaluation pitfalls—including “safety washing” and “model sandbagging”; formally define safety-critical capabilities and hazardous propensities; provide practitioners with actionable assessment guidelines; establish decision-support interfaces for regulators; and uncover fundamental research gaps in scalable, interpretable, and governance-aligned safety evaluation.
Quality engineers lack systematic degradation modeling methodologies, hindering the accuracy and practical implementation of reliability assessment. Method: This study establishes an industrially oriented degradation analysis framework that unifies diverse degradation data sources—including repeated measurements and accelerated destructive testing—and integrates path models (e.g., general path models) with stochastic process models (e.g., Wiener processes), augmented by Bayesian and likelihood-based statistical inference techniques. A standardized modeling workflow and lifetime prediction toolkit are implemented in R/Python. Contribution/Results: The framework bridges the gap between theoretical degradation modeling and engineering practice, significantly improving the accuracy and reproducibility of reliability predictions for complex systems. It delivers an actionable guideline and open-source software support for industry-standardized deployment, enabling robust, traceable, and scalable reliability engineering.
Machine learning models frequently suffer unexpected failures in real-world deployment, hindering practical adoption. Method: This paper introduces, for the first time, an orthogonal dichotomy framework distinguishing reliability from robustness, formally characterizing model failure mechanisms from first principles and systematically mapping them to engineering practices and real-world deployment scenarios. Our approach integrates probabilistic modeling, uncertainty quantification, adversarial robustness analysis, distributional shift detection, and system-level fault tree analysis—bridging theoretical insights with industrial-grade diagnostic tools and canonical failure case studies. Contribution/Results: We deliver an actionable failure attribution guide comprising rigorous theoretical foundations, an open-source toolchain, and cross-domain application exemplars. The framework significantly enhances model trustworthiness, debuggability, and deployment success rates.
Current cyber-physical systems (CPS) in vehicular environments lack quantitative, experimentally grounded methods for assessing network resilience. Method: This study constructs an experimental testbed replicating real-world truck operational conditions and conducts multiple rounds of malware injection attacks, simultaneously collecting network- and physical-layer data on resistance and recovery behaviors. Contribution/Results: We introduce the novel concept of “bonware” to holistically characterize both cybersecurity defense capability and physical resilience, formalized via an analytically tractable mathematical model. We further define and extract experimentally identifiable, quantitative resilience metrics—termed elastic features—for the first time. Sensitivity analysis confirms these metrics exhibit significant discriminability with respect to attack intensity, defensive strategies, and physical redundancy. This work bridges a critical gap by advancing vehicular CPS resilience from qualitative description to quantifiable, comparable, and optimizable measurement.
This paper addresses the reliability of calibration evaluation for machine learning models, identifying systematic biases in the widely used Expected Calibration Error (ECE) under distributional shift and varying binning strategies. Methodologically, it clarifies the logical hierarchy among multi-level calibration definitions, and systematically exposes ECE’s limitations through visualization, binning-based statistical analysis, and theoretical derivation—demonstrating its failure to satisfy key requirements of robustness and consistency in calibration assessment. Building on this critique, the paper introduces and explicates emerging calibration paradigms—including distribution-level and instance-level calibration—alongside their corresponding evaluation methodologies, thereby constructing a rigorous, interpretable, and practice-oriented calibration knowledge framework. The results equip researchers with principled guidance for selecting appropriate evaluation metrics and advance calibration assessment from ad hoc, heuristic practices toward standardization and formalization.
This work addresses the limitation of existing safety-critical systems, which typically evaluate only predictive accuracy while lacking rigorous validation of the overall calibration of predicted probability distributions. To bridge this gap, the authors propose a modular calibration testing framework that decouples the calibration process into four interchangeable components: data model, scoring rule, hypothesis formulation, and statistical test procedure. Built upon formal statistical hypothesis testing, the framework provides a single accept/reject decision for the entire predictive distribution. Crucially, it rejects only overly confident predictions while tolerating reasonable deviations, thereby balancing practicality with flexibility. Empirical evaluations on weather forecasting and robotic pose estimation tasks demonstrate that the framework effectively supports reliable deployment in safety-critical applications.
Structural reliability analysis heavily relies on specialized expertise, which limits its broader engineering application. This work proposes a multi-agent large language model framework that, for the first time, integrates a fine-tuned Method Planner with a multi-agent architecture to automate the entire component-level reliability analysis pipeline—from natural language problem descriptions through modeling, method planning, code generation, execution, and result interpretation—while incorporating human verification at critical decision points. By delegating computations to validated deterministic solvers rather than relying on the LLM to generate numerical results directly, the system significantly enhances reproducibility and mitigates hallucination. Experimental results demonstrate that the proposed approach lowers the expertise barrier while preserving the accuracy and trustworthiness of the computational outcomes.
Low-granularity operational data can lead to overly optimistic assessments of autonomous driving software reliability, thereby undermining the credibility of safety certification. This work proposes a systematic approach based on Conservative Bayesian Inference (CBI) to quantify, for the first time, the adverse impact of insufficient data fidelity on the robustness of reliability claims. By integrating statistical robustness analysis with software reliability modeling, the study demonstrates that even conservative inference strategies may yield misleading conclusions when applied to low-fidelity data. The paper establishes the first conservative estimation framework that explicitly accounts for the influence of data granularity on reliability assessment, highlighting the critical importance of high-fidelity operational data in safety certification of autonomous driving systems.
This work addresses the challenge that existing model evaluation methods often fail to reliably assess estimator quality in low-variance settings due to confounding between bias and variance or excessive sensitivity of statistical tests. To overcome this limitation, the authors propose a fault-tolerant evaluation framework that unifies bias and variance modeling through an adjustable tolerance parameter ε, enabling robust assessment of sample-efficient performance estimators within practically acceptable error margins. The framework integrates bias-variance analysis, fault-tolerant evaluation theory, and an adaptive ε-optimization algorithm, making it particularly well-suited for scenarios with low annotation costs. Experimental results demonstrate that the proposed approach provides a more comprehensive and reliable characterization of estimator behavior, significantly enhancing both the practical utility and stability of performance evaluation.
This study addresses the limitations of traditional process capability indices, such as Cpk, which rely on deterministic thresholds under finite sample sizes and neglect estimation uncertainty, often leading to misclassification in critical regions. To overcome this, the work reframes capability assessment as a decision-risk calibration problem and introduces a hybrid framework that integrates a statistical baseline with data-driven residual learning—marking the first incorporation of uncertainty quantification into process capability evaluation. The approach leverages an interpretable baseline to model prior structural assumptions while employing a residual network to capture deviations due to non-normality, measurement error, and small-sample bias. Decision-risk calibration is achieved through nested Monte Carlo simulation. Experimental results demonstrate that the proposed framework significantly improves calibration accuracy and stability in critical regions, remains robust under leakage-free evaluation, and is readily deployable within existing industrial systems.