robustness testing

Designs and implements test suites, evaluation frameworks, benchmarks, metrics and validation procedures to measure how models behave under noise, perturbations, adversarial inputs and distribution shifts. Analyzes robustness with statistical and empirical methods and builds engineering and optimization techniques, checks, reports and tooling to validate, enhance and continuously monitor model robustness.

robustnesstesting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.66
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$220K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of pinpointing and tracing error sources and propagation pathways within composite AI systems comprising multiple neural network components, a task that existing robustness testing methods struggle to accomplish. To this end, the paper proposes a modular robustness testing framework that enables fine-grained fault attribution through statistical perturbation injection, component-level error tracking, and cross-module propagation inference. By moving beyond conventional end-to-end evaluation paradigms, the approach supports architecture- and modality-agnostic analysis, offering a generalized methodology for dissecting system-level robustness. The framework’s efficacy is demonstrated in a railway track inspection system, where it reveals nuanced robustness characteristics that surpass the diagnostic granularity of standard evaluation metrics.

compound AI systemserror attributionmodular analysis

Evaluating the Evaluators: Trust in Adversarial Robustness Tests

Jul 04, 2025
AE
Antonio Emanuele Cinà
🏛️ University of Genoa | Ca’ Foscari University of Venice

Inconsistent and unreliable adversarial robustness evaluations arise from model mismatch, non-verifiable implementations, and unequal computational budgets. To address these issues, this paper introduces AttackBench—a standardized benchmarking framework. AttackBench unifies evaluation using gradient-based attacks, a curated set of standard models, and fully reproducible implementations; it further proposes a novel optimality-based metric and strictly controls experimental conditions to ensure fair comparisons. The framework enables trustworthy ranking of mainstream attack methods, systematically identifies sources of bias in existing evaluations, and significantly improves the reproducibility and credibility of robustness verification. Its modular architecture supports continuous extension and benchmark updates, providing a reliable, open evaluation infrastructure for adversarial robustness research.

Flawed testing protocols leading to misleading robustness claimsInconsistent evaluation of adversarial evasion attacks methodsLack of standardized conditions for assessing gradient-based attacks

Measuring training variability from stochastic optimization using robust nonparametric testing

Jun 12, 2024
SB
Sinjini Banerjee
🏛️ Rutgers University | Pacific Northwest National Lab

Deep neural network training suffers from high sensitivity to random seeds due to stochastic optimization, hindering reliable assessment of true generalization performance. To address this, we propose a robust nonparametric hypothesis testing framework. Its core innovation is a novel model similarity metric—the α-truncation level—which quantifies training variability and determines the minimum number of independent training runs required for stable ensembling. Unlike conventional metrics such as accuracy or expected calibration error (ECE), the α-truncation level does not rely on modeling the null distribution and is inherently sensitive to training instability. Experiments demonstrate that it detects training uncertainty earlier and more consistently than validation accuracy, churn, and ECE. Moreover, in transfer learning settings, it effectively guides random seed selection, significantly improving the reliability of performance evaluation.

Determining optimal training runs for reliable ensemble performanceDeveloping robust hypothesis testing for model similarity assessmentMeasuring variability in deep neural network training outcomes

To address the degradation of model robustness post-deployment caused by hardware/software environment shifts, this paper introduces Prom, an open-source framework that pioneers dynamic misprediction detection and lightweight feedback-driven adaptive repair at deployment time. The method integrates statistical significance testing, uncertainty quantification, and confidence calibration—enabling accuracy recovery without full retraining. Instead, it leverages an online feedback loop to incrementally annotate and learn from ≤5% of samples. Evaluated across 13 models and five code analysis and optimization tasks, Prom achieves an average misprediction identification rate of 96% (up to 100%), significantly enhancing cross-platform generalization and robustness against diverse hardware configurations and code patterns.

Model RobustnessPrediction StabilityReliability under Hardware/Software Changes

Latest Papers

What's happening recently
View more

This work addresses the challenge of reliably detecting performance degradation in large language models caused by optimization techniques such as quantization, where observed drops in accuracy may stem from genuine model deterioration or mere evaluation noise. To this end, the authors propose a statistical hypothesis testing framework based on McNemar’s test, which introduces sample-level paired comparisons for the first time in the context of LLM degradation analysis, thereby overcoming the limited sensitivity of conventional task-level aggregation. Integrated with multi-benchmark accuracy aggregation and the LM Evaluation Harness, the method effectively controls false positive rates and reliably identifies performance degradations as small as 0.3%. Empirical results demonstrate that the approach accurately flags models exhibiting true degradation while producing no false alarms for theoretically lossless optimizations.

accuracy evaluationLLM optimizationmodel degradation

This study addresses the critical issue of model instability in software engineering optimization, which leads to substantial variability across repeated experiments and undermines both credibility and practical utility. Rather than treating instability as mere random noise, this work conceptualizes it as a quantifiable and manageable property that should be integrated into standard evaluation frameworks. By systematically modulating label usage, model complexity, and partition scoring strategies—combined with multi-objective optimization, causal intervention, data locality analysis, and model calibration—the proposed approach significantly enhances result consistency. Empirical evaluation demonstrates that the optimized configuration reduces the standard deviation of error by 22% on average and outperforms default settings in 119 out of 127 datasets, achieving a 4.8-fold improvement in result consistency.

model instabilitymulti-objective optimizationreproducibility

Existing software modeling datasets are often ad hoc constructions lacking rigorous quality assurance, leading to research findings that are difficult to reproduce, compare, and prone to bias. This work proposes the first benchmarking framework specifically designed for model-driven engineering, treating datasets themselves as first-class evaluation targets. By defining clear metrics for quality, representativeness, and task suitability, the framework establishes a unified platform that enables automated analysis of modeling datasets across multiple languages and formats. For the first time, this approach facilitates systematic evaluation of modeling datasets, substantially enhancing the reproducibility, fairness, and scientific rigor of research in the field.

benchmarkingdataset qualitymodel datasets

This study investigates the relationship between the robustness of neural networks under random input perturbations and their prediction accuracy, measured by mean squared error (MSE). To address this, the work proposes an efficient, computable black-box robustness metric that, without requiring access to internal model architecture, provides a high-probability upper bound on the network’s MSE over an entire dataset under a given perturbation. The method innovatively introduces robustness curves, enabling systematic comparison and analysis of robustness across different datasets. Experimental evaluations on multiple real-world datasets demonstrate that the proposed approach accurately quantifies and effectively captures a model’s sensitivity to input noise, offering a practical tool for assessing robustness in diverse settings.

input perturbationsmean squared errorneural networks

Existing Simulink model checkers often produce verification results inconsistent with simulation outcomes due to the absence of bit-precise formal semantics for modeling elements and numerical behaviors, undermining their reliability. This work proposes the first bit-precise conformance testing methodology tailored for Simulink model checkers. By formally specifying the semantics of fundamental blocks, constructing a test suite covering ten block categories, and integrating SMT solving within an automated framework, the approach systematically evaluates behavioral alignment across tools. Experimental results demonstrate that the method effectively uncovers inconsistencies: while the third-party checker SmtMC passes all tests, Simulink Design Verifier exhibits only 94–96% conformance with the simulator, and its agreement with other checkers drops further to 80–90%. The framework also precisely identifies the root causes of these discrepancies.

bit-precise conformancecyber-physical systemsformal verification

Hot Scholars

EC

Enhong Chen

University of Science and Technology of China
data miningrecommender systemmachine learning
HS

Hwanjun Song

Assistant Professor, KAIST
LLMTrustworthy AIHuman-AI AlignmentData-centric AI
DT

Dacheng Tao

Nanyang Technological University
artificial intelligencemachine learningcomputer visionimage processing
XM

Xingjun Ma

Fudan University
Trustworthy AIMultimodal AIGenerative AIEmbodied AI
CX

Cihang Xie

Assistant Professor, University of California, Santa Cruz
Computer VisionMachine Learning