robustness

Designs, builds, and evaluates models, algorithms, and systems that maintain acceptable performance when inputs, environments, or operational conditions deviate from training assumptions — e.g., under distribution shift, noise, adversarial perturbations, missing or corrupted data, or implementation variability. Work includes creating robustness tests and benchmarks, devising defenses or resilient architectures, deriving worst‑case or probabilistic guarantees and failure‑mode analyses, and measuring sensitivity and stability.

robustness

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
1.56
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$213K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the coupled challenges of model pruning, adversarial perturbations, and stuck-at-zero hardware weight faults in deep neural networks deployed on resource-constrained neuromorphic hardware, where their joint impact remains poorly understood. Focusing on a three-layer CNN trained on MNIST, the work presents a systematic multidimensional evaluation integrating pruning, adversarial training, and hardware fault injection. It reveals, for the first time, that while adversarial training enhances robustness against attacks such as FGSM, it substantially increases sensitivity to stuck-at-zero faults. In contrast, pruning exhibits limited influence on fault tolerance and demonstrates consistent performance across varying fault rates and attack intensities, challenging prevailing assumptions in the field.

adversarial robustnessfault tolerancehardware faults

Evaluating the Evaluators: Trust in Adversarial Robustness Tests

Jul 04, 2025
AE
Antonio Emanuele Cinà
🏛️ University of Genoa | Ca’ Foscari University of Venice

Inconsistent and unreliable adversarial robustness evaluations arise from model mismatch, non-verifiable implementations, and unequal computational budgets. To address these issues, this paper introduces AttackBench—a standardized benchmarking framework. AttackBench unifies evaluation using gradient-based attacks, a curated set of standard models, and fully reproducible implementations; it further proposes a novel optimality-based metric and strictly controls experimental conditions to ensure fair comparisons. The framework enables trustworthy ranking of mainstream attack methods, systematically identifies sources of bias in existing evaluations, and significantly improves the reproducibility and credibility of robustness verification. Its modular architecture supports continuous extension and benchmark updates, providing a reliable, open evaluation infrastructure for adversarial robustness research.

Flawed testing protocols leading to misleading robustness claimsInconsistent evaluation of adversarial evasion attacks methodsLack of standardized conditions for assessing gradient-based attacks

Machine learning models frequently suffer unexpected failures in real-world deployment, hindering practical adoption. Method: This paper introduces, for the first time, an orthogonal dichotomy framework distinguishing reliability from robustness, formally characterizing model failure mechanisms from first principles and systematically mapping them to engineering practices and real-world deployment scenarios. Our approach integrates probabilistic modeling, uncertainty quantification, adversarial robustness analysis, distributional shift detection, and system-level fault tree analysis—bridging theoretical insights with industrial-grade diagnostic tools and canonical failure case studies. Contribution/Results: We deliver an actionable failure attribution guide comprising rigorous theoretical foundations, an open-source toolchain, and cross-domain application exemplars. The framework significantly enhances model trustworthiness, debuggability, and deployment success rates.

Differentiating failure causes: reliability vs robustnessProviding practical techniques for ML model reliabilityUnderstanding unexpected failures in ML models

To address the degradation of model robustness post-deployment caused by hardware/software environment shifts, this paper introduces Prom, an open-source framework that pioneers dynamic misprediction detection and lightweight feedback-driven adaptive repair at deployment time. The method integrates statistical significance testing, uncertainty quantification, and confidence calibration—enabling accuracy recovery without full retraining. Instead, it leverages an online feedback loop to incrementally annotate and learn from ≤5% of samples. Evaluated across 13 models and five code analysis and optimization tasks, Prom achieves an average misprediction identification rate of 96% (up to 100%), significantly enhancing cross-platform generalization and robustness against diverse hardware configurations and code patterns.

Model RobustnessPrediction StabilityReliability under Hardware/Software Changes

Latest Papers

What's happening recently
View more

Current evaluation datasets struggle to accurately estimate the risk of rare failures that machine learning models may encounter in deployment. This work proposes an extrapolation method for failure rates grounded in extreme value theory, leveraging the top-k largest failure scores observed in the evaluation set to predict failure rates at deployment scale. To address the inherent safety bias and the tendency of existing extrapolation estimators to overlook high-risk failure modes, the approach incorporates a predictability-aware loss function during fine-tuning. Experiments on the Password Game and GridWorld benchmarks demonstrate that the proposed method substantially reduces prediction error while preserving primary task performance, achieving safety levels comparable to those of supervised baselines.

deployment-scale failure rateevaluation set limitationfailure prediction

This study investigates the relationship between the robustness of neural networks under random input perturbations and their prediction accuracy, measured by mean squared error (MSE). To address this, the work proposes an efficient, computable black-box robustness metric that, without requiring access to internal model architecture, provides a high-probability upper bound on the network’s MSE over an entire dataset under a given perturbation. The method innovatively introduces robustness curves, enabling systematic comparison and analysis of robustness across different datasets. Experimental evaluations on multiple real-world datasets demonstrate that the proposed approach accurately quantifies and effectively captures a model’s sensitivity to input noise, offering a practical tool for assessing robustness in diverse settings.

input perturbationsmean squared errorneural networks

This work addresses a critical gap in existing defenses against malicious fine-tuning, which are typically effective only against predefined attacks and lack robustness against adaptive adversaries. We propose the first unified adaptive attack framework that integrates adversarial fine-tuning, path analysis, and attack construction techniques to systematically evaluate 15 state-of-the-art defense mechanisms. Our comprehensive assessment demonstrates that all evaluated defenses can be successfully circumvented, as they merely obscure or misdirect the activation pathways of harmful behaviors without eliminating the underlying capabilities. These findings expose fundamental limitations in current alignment-based defenses and establish a reliable benchmark and clear direction for future research on robustly securing language models against adaptive threats.

adaptive adversariesdefense mechanismsmalicious fine-tuning