Score
Design and implement evaluation methods and estimators that quantify a model or system's robustness by estimating probabilities of failure or performance degradation under stochastic or adversarial perturbations. Produce statistical summaries and tests—confidence intervals, significance tests, and comparative risk metrics—to compare reliability across models and support deployment decisions.
Quantifying how input uncertainty propagates to model outputs remains a fundamental challenge in computational modeling. Method: This study systematically reviews and empirically compares prominent global and local sensitivity analysis (SA) techniques—including Sobol’, FAST, Morris screening, and local derivative-based methods—implemented via standard software packages, supporting both probabilistic modeling and distribution-free settings. Contribution/Results: We propose a practical decision framework that guides method selection based on problem characteristics, analytical objectives, and resource constraints—rejecting the notion of a universally “optimal” SA method and thereby addressing a critical gap in methodological implementation guidance. A reusable, open-source toolkit is developed to enhance the reliability and interpretability of uncertainty attribution. The framework and tools have been validated across multiple engineering and policy modeling applications, demonstrating robustness and scalability in real-world contexts.
This work addresses the challenge that existing model evaluation methods often fail to reliably assess estimator quality in low-variance settings due to confounding between bias and variance or excessive sensitivity of statistical tests. To overcome this limitation, the authors propose a fault-tolerant evaluation framework that unifies bias and variance modeling through an adjustable tolerance parameter ε, enabling robust assessment of sample-efficient performance estimators within practically acceptable error margins. The framework integrates bias-variance analysis, fault-tolerant evaluation theory, and an adaptive ε-optimization algorithm, making it particularly well-suited for scenarios with low annotation costs. Experimental results demonstrate that the proposed approach provides a more comprehensive and reliable characterization of estimator behavior, significantly enhancing both the practical utility and stability of performance evaluation.
To address the longstanding trade-off between computational cost and accuracy in pre-deployment robustness assessment for safety-critical applications, this paper proposes a hypothesis-testing-based quantitative evaluation framework. Our core contribution is the introduction of “tower robustness”—a novel metric that, for the first time, incorporates statistical hypothesis testing into probabilistic modeling of deep learning robustness, enabling rigorous, verifiable quantification of model output stability under input perturbations. By integrating probabilistic modeling with comparative analysis, the framework systematically restructures the evaluation pipeline, achieving both theoretical soundness and substantial efficiency gains. Extensive experiments across large-scale benchmarks demonstrate that our approach improves assessment accuracy by 12.7% on average and reduces runtime by 43.5% compared to state-of-the-art baselines. This work establishes a new paradigm for pre-deployment risk analysis of high-assurance AI systems—one that is both practically deployable and inherently interpretable.
Machine learning models frequently suffer unexpected failures in real-world deployment, hindering practical adoption. Method: This paper introduces, for the first time, an orthogonal dichotomy framework distinguishing reliability from robustness, formally characterizing model failure mechanisms from first principles and systematically mapping them to engineering practices and real-world deployment scenarios. Our approach integrates probabilistic modeling, uncertainty quantification, adversarial robustness analysis, distributional shift detection, and system-level fault tree analysis—bridging theoretical insights with industrial-grade diagnostic tools and canonical failure case studies. Contribution/Results: We deliver an actionable failure attribution guide comprising rigorous theoretical foundations, an open-source toolchain, and cross-domain application exemplars. The framework significantly enhances model trustworthiness, debuggability, and deployment success rates.
This work addresses the challenge of estimating small failure probabilities under stochastic inputs in computationally expensive deterministic simulations. We propose a two-stage adaptive budget allocation framework: in Stage I, a Gaussian process surrogate is sequentially trained using a contour-localization strategy; in Stage II, remaining simulation budget is greedily allocated to critical regions—guided by classification entropy—to perform high-fidelity evaluations. A hybrid Monte Carlo estimator is then constructed by integrating surrogate predictions with observed high-fidelity responses. Our method introduces the first “exploration–exploitation decoupled” budget allocation paradigm, overcoming reliability limitations inherent in pure surrogate-based Monte Carlo and importance sampling. Experiments across multiple benchmark functions and an airfoil flow simulation demonstrate that the approach achieves significantly improved accuracy and robustness using only several hundred high-fidelity evaluations.
Current cyber-physical systems (CPS) in vehicular environments lack quantitative, experimentally grounded methods for assessing network resilience. Method: This study constructs an experimental testbed replicating real-world truck operational conditions and conducts multiple rounds of malware injection attacks, simultaneously collecting network- and physical-layer data on resistance and recovery behaviors. Contribution/Results: We introduce the novel concept of “bonware” to holistically characterize both cybersecurity defense capability and physical resilience, formalized via an analytically tractable mathematical model. We further define and extract experimentally identifiable, quantitative resilience metrics—termed elastic features—for the first time. Sensitivity analysis confirms these metrics exhibit significant discriminability with respect to attack intensity, defensive strategies, and physical redundancy. This work bridges a critical gap by advancing vehicular CPS resilience from qualitative description to quantifiable, comparable, and optimizable measurement.
Current evaluation datasets struggle to accurately estimate the risk of rare failures that machine learning models may encounter in deployment. This work proposes an extrapolation method for failure rates grounded in extreme value theory, leveraging the top-k largest failure scores observed in the evaluation set to predict failure rates at deployment scale. To address the inherent safety bias and the tendency of existing extrapolation estimators to overlook high-risk failure modes, the approach incorporates a predictability-aware loss function during fine-tuning. Experiments on the Password Game and GridWorld benchmarks demonstrate that the proposed method substantially reduces prediction error while preserving primary task performance, achieving safety levels comparable to those of supervised baselines.
This study investigates the relationship between the robustness of neural networks under random input perturbations and their prediction accuracy, measured by mean squared error (MSE). To address this, the work proposes an efficient, computable black-box robustness metric that, without requiring access to internal model architecture, provides a high-probability upper bound on the network’s MSE over an entire dataset under a given perturbation. The method innovatively introduces robustness curves, enabling systematic comparison and analysis of robustness across different datasets. Experimental evaluations on multiple real-world datasets demonstrate that the proposed approach accurately quantifies and effectively captures a model’s sensitivity to input noise, offering a practical tool for assessing robustness in diverse settings.
This study addresses the pronounced amplification of estimation error in process capability indices under limited sample sizes when nonlinearly transformed into tail-risk metrics such as defect probability or parts per million (PPM), which undermines decision stability. For the first time, it elucidates the nonlinear propagation mechanism of uncertainty between capability index space and tail-risk space, offering a unified explanation for the unreliability of quality decisions in small-sample settings. The work quantitatively links required sample size to decision reliability and, through Monte Carlo simulations, industrial data validation, and statistical inference, systematically models the nonlinear relationship between process capability indices and defect risk. It further clarifies how distributional assumptions critically influence risk estimation, thereby establishing a theoretical foundation and practical guidance for reliability-aware quality decision-making.