Score
Designs and implements test suites, evaluation frameworks, benchmarks, metrics and validation procedures to measure how models behave under noise, perturbations, adversarial inputs and distribution shifts. Analyzes robustness with statistical and empirical methods and builds engineering and optimization techniques, checks, reports and tooling to validate, enhance and continuously monitor model robustness.
This work addresses the challenge of pinpointing and tracing error sources and propagation pathways within composite AI systems comprising multiple neural network components, a task that existing robustness testing methods struggle to accomplish. To this end, the paper proposes a modular robustness testing framework that enables fine-grained fault attribution through statistical perturbation injection, component-level error tracking, and cross-module propagation inference. By moving beyond conventional end-to-end evaluation paradigms, the approach supports architecture- and modality-agnostic analysis, offering a generalized methodology for dissecting system-level robustness. The framework’s efficacy is demonstrated in a railway track inspection system, where it reveals nuanced robustness characteristics that surpass the diagnostic granularity of standard evaluation metrics.
Inconsistent and unreliable adversarial robustness evaluations arise from model mismatch, non-verifiable implementations, and unequal computational budgets. To address these issues, this paper introduces AttackBench—a standardized benchmarking framework. AttackBench unifies evaluation using gradient-based attacks, a curated set of standard models, and fully reproducible implementations; it further proposes a novel optimality-based metric and strictly controls experimental conditions to ensure fair comparisons. The framework enables trustworthy ranking of mainstream attack methods, systematically identifies sources of bias in existing evaluations, and significantly improves the reproducibility and credibility of robustness verification. Its modular architecture supports continuous extension and benchmark updates, providing a reliable, open evaluation infrastructure for adversarial robustness research.
Deep neural network training suffers from high sensitivity to random seeds due to stochastic optimization, hindering reliable assessment of true generalization performance. To address this, we propose a robust nonparametric hypothesis testing framework. Its core innovation is a novel model similarity metric—the α-truncation level—which quantifies training variability and determines the minimum number of independent training runs required for stable ensembling. Unlike conventional metrics such as accuracy or expected calibration error (ECE), the α-truncation level does not rely on modeling the null distribution and is inherently sensitive to training instability. Experiments demonstrate that it detects training uncertainty earlier and more consistently than validation accuracy, churn, and ECE. Moreover, in transfer learning settings, it effectively guides random seed selection, significantly improving the reliability of performance evaluation.
To address the degradation of model robustness post-deployment caused by hardware/software environment shifts, this paper introduces Prom, an open-source framework that pioneers dynamic misprediction detection and lightweight feedback-driven adaptive repair at deployment time. The method integrates statistical significance testing, uncertainty quantification, and confidence calibration—enabling accuracy recovery without full retraining. Instead, it leverages an online feedback loop to incrementally annotate and learn from ≤5% of samples. Evaluated across 13 models and five code analysis and optimization tasks, Prom achieves an average misprediction identification rate of 96% (up to 100%), significantly enhancing cross-platform generalization and robustness against diverse hardware configurations and code patterns.
This work addresses the challenge of reliably detecting performance degradation in large language models caused by optimization techniques such as quantization, where observed drops in accuracy may stem from genuine model deterioration or mere evaluation noise. To this end, the authors propose a statistical hypothesis testing framework based on McNemar’s test, which introduces sample-level paired comparisons for the first time in the context of LLM degradation analysis, thereby overcoming the limited sensitivity of conventional task-level aggregation. Integrated with multi-benchmark accuracy aggregation and the LM Evaluation Harness, the method effectively controls false positive rates and reliably identifies performance degradations as small as 0.3%. Empirical results demonstrate that the approach accurately flags models exhibiting true degradation while producing no false alarms for theoretically lossless optimizations.
This study addresses the critical issue of model instability in software engineering optimization, which leads to substantial variability across repeated experiments and undermines both credibility and practical utility. Rather than treating instability as mere random noise, this work conceptualizes it as a quantifiable and manageable property that should be integrated into standard evaluation frameworks. By systematically modulating label usage, model complexity, and partition scoring strategies—combined with multi-objective optimization, causal intervention, data locality analysis, and model calibration—the proposed approach significantly enhances result consistency. Empirical evaluation demonstrates that the optimized configuration reduces the standard deviation of error by 22% on average and outperforms default settings in 119 out of 127 datasets, achieving a 4.8-fold improvement in result consistency.
Existing software modeling datasets are often ad hoc constructions lacking rigorous quality assurance, leading to research findings that are difficult to reproduce, compare, and prone to bias. This work proposes the first benchmarking framework specifically designed for model-driven engineering, treating datasets themselves as first-class evaluation targets. By defining clear metrics for quality, representativeness, and task suitability, the framework establishes a unified platform that enables automated analysis of modeling datasets across multiple languages and formats. For the first time, this approach facilitates systematic evaluation of modeling datasets, substantially enhancing the reproducibility, fairness, and scientific rigor of research in the field.
This study investigates the relationship between the robustness of neural networks under random input perturbations and their prediction accuracy, measured by mean squared error (MSE). To address this, the work proposes an efficient, computable black-box robustness metric that, without requiring access to internal model architecture, provides a high-probability upper bound on the network’s MSE over an entire dataset under a given perturbation. The method innovatively introduces robustness curves, enabling systematic comparison and analysis of robustness across different datasets. Experimental evaluations on multiple real-world datasets demonstrate that the proposed approach accurately quantifies and effectively captures a model’s sensitivity to input noise, offering a practical tool for assessing robustness in diverse settings.
Existing Simulink model checkers often produce verification results inconsistent with simulation outcomes due to the absence of bit-precise formal semantics for modeling elements and numerical behaviors, undermining their reliability. This work proposes the first bit-precise conformance testing methodology tailored for Simulink model checkers. By formally specifying the semantics of fundamental blocks, constructing a test suite covering ten block categories, and integrating SMT solving within an automated framework, the approach systematically evaluates behavioral alignment across tools. Experimental results demonstrate that the method effectively uncovers inconsistencies: while the third-party checker SmtMC passes all tests, Simulink Design Verifier exhibits only 94–96% conformance with the simulator, and its agreement with other checkers drops further to 80–90%. The framework also precisely identifies the root causes of these discrepancies.