Score
Designs, builds, and evaluates models, algorithms, and systems that maintain acceptable performance when inputs, environments, or operational conditions deviate from training assumptions — e.g., under distribution shift, noise, adversarial perturbations, missing or corrupted data, or implementation variability. Work includes creating robustness tests and benchmarks, devising defenses or resilient architectures, deriving worst‑case or probabilistic guarantees and failure‑mode analyses, and measuring sensitivity and stability.
本文探讨了在数据不完美条件下的机器学习挑战,并通过信息损失、经验风险偏差等机制组织代表性方法,如重建生成、再平衡与表示校准等。
This study addresses the coupled challenges of model pruning, adversarial perturbations, and stuck-at-zero hardware weight faults in deep neural networks deployed on resource-constrained neuromorphic hardware, where their joint impact remains poorly understood. Focusing on a three-layer CNN trained on MNIST, the work presents a systematic multidimensional evaluation integrating pruning, adversarial training, and hardware fault injection. It reveals, for the first time, that while adversarial training enhances robustness against attacks such as FGSM, it substantially increases sensitivity to stuck-at-zero faults. In contrast, pruning exhibits limited influence on fault tolerance and demonstrates consistent performance across varying fault rates and attack intensities, challenging prevailing assumptions in the field.
Inconsistent and unreliable adversarial robustness evaluations arise from model mismatch, non-verifiable implementations, and unequal computational budgets. To address these issues, this paper introduces AttackBench—a standardized benchmarking framework. AttackBench unifies evaluation using gradient-based attacks, a curated set of standard models, and fully reproducible implementations; it further proposes a novel optimality-based metric and strictly controls experimental conditions to ensure fair comparisons. The framework enables trustworthy ranking of mainstream attack methods, systematically identifies sources of bias in existing evaluations, and significantly improves the reproducibility and credibility of robustness verification. Its modular architecture supports continuous extension and benchmark updates, providing a reliable, open evaluation infrastructure for adversarial robustness research.
Machine learning models frequently suffer unexpected failures in real-world deployment, hindering practical adoption. Method: This paper introduces, for the first time, an orthogonal dichotomy framework distinguishing reliability from robustness, formally characterizing model failure mechanisms from first principles and systematically mapping them to engineering practices and real-world deployment scenarios. Our approach integrates probabilistic modeling, uncertainty quantification, adversarial robustness analysis, distributional shift detection, and system-level fault tree analysis—bridging theoretical insights with industrial-grade diagnostic tools and canonical failure case studies. Contribution/Results: We deliver an actionable failure attribution guide comprising rigorous theoretical foundations, an open-source toolchain, and cross-domain application exemplars. The framework significantly enhances model trustworthiness, debuggability, and deployment success rates.
To address the degradation of model robustness post-deployment caused by hardware/software environment shifts, this paper introduces Prom, an open-source framework that pioneers dynamic misprediction detection and lightweight feedback-driven adaptive repair at deployment time. The method integrates statistical significance testing, uncertainty quantification, and confidence calibration—enabling accuracy recovery without full retraining. Instead, it leverages an online feedback loop to incrementally annotate and learn from ≤5% of samples. Evaluated across 13 models and five code analysis and optimization tasks, Prom achieves an average misprediction identification rate of 96% (up to 100%), significantly enhancing cross-platform generalization and robustness against diverse hardware configurations and code patterns.
研究通过GitHub挖掘分析了28个开源机器学习鲁棒性评估工具的维护和持续性问题,发现活跃度不均,强调需将这些工具视为不断发展的软件系统。
本文通过将韧性理论应用于ICS异常检测,解决了系统级韧性认证问题,提出了一种新的评估方法,并在BATADAL水分配系统上进行了验证。
Current evaluation datasets struggle to accurately estimate the risk of rare failures that machine learning models may encounter in deployment. This work proposes an extrapolation method for failure rates grounded in extreme value theory, leveraging the top-k largest failure scores observed in the evaluation set to predict failure rates at deployment scale. To address the inherent safety bias and the tendency of existing extrapolation estimators to overlook high-risk failure modes, the approach incorporates a predictability-aware loss function during fine-tuning. Experiments on the Password Game and GridWorld benchmarks demonstrate that the proposed method substantially reduces prediction error while preserving primary task performance, achieving safety levels comparable to those of supervised baselines.
This study investigates the relationship between the robustness of neural networks under random input perturbations and their prediction accuracy, measured by mean squared error (MSE). To address this, the work proposes an efficient, computable black-box robustness metric that, without requiring access to internal model architecture, provides a high-probability upper bound on the network’s MSE over an entire dataset under a given perturbation. The method innovatively introduces robustness curves, enabling systematic comparison and analysis of robustness across different datasets. Experimental evaluations on multiple real-world datasets demonstrate that the proposed approach accurately quantifies and effectively captures a model’s sensitivity to input noise, offering a practical tool for assessing robustness in diverse settings.
This work addresses a critical gap in existing defenses against malicious fine-tuning, which are typically effective only against predefined attacks and lack robustness against adaptive adversaries. We propose the first unified adaptive attack framework that integrates adversarial fine-tuning, path analysis, and attack construction techniques to systematically evaluate 15 state-of-the-art defense mechanisms. Our comprehensive assessment demonstrates that all evaluated defenses can be successfully circumvented, as they merely obscure or misdirect the activation pathways of harmful behaviors without eliminating the underlying capabilities. These findings expose fundamental limitations in current alignment-based defenses and establish a reliable benchmark and clear direction for future research on robustly securing language models against adaptive threats.