pre-registered robustness evaluation

Designs and implements pre-registered evaluation protocols, metrics, and analysis plans to quantify a system's empirical robustness to noise, corruptions, and other perturbations; this includes specifying perturbation sets, comparative degradation hypotheses, evaluation metrics (accuracy and robustness metrics), and pre-specified statistical analyses. Builds reproducible assessment pipelines and reporting procedures that apply production-relevant perturbations, measure performance across conditions and controllers, and disclose outcomes according to the pre-registration.

pre-registeredrobustnessevaluation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.2
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$224K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Get Global Guarantees: On the Probabilistic Nature of Perturbation Robustness

Aug 26, 2025
WM
Wenchuan Mu
🏛️ Singapore University of Technology and Design

To address the longstanding trade-off between computational cost and accuracy in pre-deployment robustness assessment for safety-critical applications, this paper proposes a hypothesis-testing-based quantitative evaluation framework. Our core contribution is the introduction of “tower robustness”—a novel metric that, for the first time, incorporates statistical hypothesis testing into probabilistic modeling of deep learning robustness, enabling rigorous, verifiable quantification of model output stability under input perturbations. By integrating probabilistic modeling with comparative analysis, the framework systematically restructures the evaluation pipeline, achieving both theoretical soundness and substantial efficiency gains. Extensive experiments across large-scale benchmarks demonstrate that our approach improves assessment accuracy by 12.7% on average and reduces runtime by 43.5% compared to state-of-the-art baselines. This work establishes a new paradigm for pre-deployment risk analysis of high-assurance AI systems—one that is both practically deployable and inherently interpretable.

Addressing computational cost and precision trade-offs in robustness assessmentEvaluating probabilistic robustness in neural networks against perturbationsProviding rigorous pre-deployment evaluation for safety-critical deep learning

This study addresses the reliability of probabilistic uncertainty quantification (UQ) in software defect prediction, particularly its ability to reflect model performance and calibration—especially in cross-project settings, where systematic validation remains lacking. Through a large-scale empirical analysis of 16 classifiers across 36 within-project and 32 cross-project datasets, the work examines the relationships between five UQ metrics and six performance measures alongside three calibration metrics. It reveals, for the first time, a strong context dependency: within-project, UQ correlates strongly with false positive rate and AUC, but these correlations substantially weaken or even reverse in cross-project scenarios. Notably, high-performing models can still exhibit severe miscalibration. These findings indicate that UQ signals are not directly transferable and must be evaluated independently relative to specific objectives, using multidimensional calibration assessments.

CalibrationCross-Project PredictionPerformance Evaluation

This work addresses the efficiency bottleneck of manual reproducibility reviews in safety-critical domains such as the Internet of Things and cyber-physical systems, which hampers research transparency and deployability. The paper presents the first systematic framework leveraging large language models (LLMs) to automate reproducibility assessment by integrating natural language understanding, code generation, sandboxed environment auto-configuration, and rule-guided flaw detection. This approach enables reproducibility scoring, automatic execution environment setup, and identification of methodological flaws. Experimental results demonstrate that the proposed method achieves over 72% accuracy in reproducibility judgment, automatically constructs executable environments for 28% of runnable artifacts, and attains F1 scores exceeding 92% across seven common categories of methodological defects, substantially enhancing both the efficiency and quality of reproducibility review.

Artifact EvaluationCPSCybersecurity

In industrial quality inspection, anomaly detection suffers from poor robustness due to high noise levels and sparse defective samples. To address this, we propose Iterative Refinement of Pseudo-labels (IRP), a self-supervised method that alternately evaluates sample credibility and removes misleading instances under feature-space consistency constraints—effectively purifying the training set dynamically without human annotations and generating high-fidelity self-supervised signals. IRP introduces the novel paradigm of “iterative data refinement,” significantly enhancing model robustness against label noise and cross-domain generalization capability. Evaluated on KSDD2 and MVTec AD benchmarks, IRP consistently outperforms existing unsupervised and self-supervised methods. Notably, under high-noise conditions, it achieves substantial improvements in detection accuracy and reduces false positive rates by over 25%.

Enhances defect detection accuracy in industrial quality control.Improves model performance by removing misleading data points.Outperforms traditional models in noisy industrial environments.

Latest Papers

What's happening recently
View more

Current AI-assisted scientific writing lacks auditable generation processes and mechanisms for accountability, undermining the verifiability of research credibility and compliance. This work proposes a novel auditing paradigm embedded directly within the production workflow, enforcing end-to-end traceability, immutability, and third-party reproducibility of AI involvement through preregistered blind-spot indicator cards, sealed execution environments, and automated gatekeeping intercepts. Core technical components include Git-sealed lineage anchoring, hash-bound provenance tracking, red-flag interception protocols, cross-model role isolation, and programmatic assembly. In experimental validation, one project was automatically terminated when preregistered confirmatory tests triggered a No-Go decision. An open-source toolkit is released to enable independent recomputation of all core audit metrics by third parties.

AI AccountabilityAuditable AIProvenance

This work addresses a critical gap in the evaluation of security vulnerability reproductions generated by large language models (LLMs) or agents: while such artifacts are often executable, they frequently lack rigorous validation confirming that they faithfully reproduce the specific CVE in question, revealing a disconnect between semantic intent and actual vulnerability signals. To remedy this, the paper introduces the first reusable evaluation protocol for security reproductions, integrating preregistered auditing, an R0/R1 environment remediation ladder, G1–G3 semantic evidence tiers, and a patch-based counterfactual oracle. Applying this framework to 104 published artifacts from 2023–2026, the study finds that 56.9% exhibit CVE identifier mismatches, only 61.1% remain functional after patching, and the oracle demonstrates limited reliability with 60% sensitivity and 45% specificity—highlighting widespread issues of false-positive patch versions and benign inputs in current approaches.

CVE verificationLLM-driven artifactsreproducibility

Hot Scholars

PY

Pin-Yu Chen

Principal Research Scientist, IBM Research AI; MIT-IBM Watson AI Lab; RPI-IBM AIRC
AI SafetyGenerative AITrustworthy Machine LearningAdversarial Machine Learning
FR

Fabio Roli

Professor, University of Genova and Cagliari, Italy
Pattern recognitionmachine learningcomputer visioncomputer security
LC

Leshem Choshen

MIT, IBM AI research
Model RecyclingEvolving Collaborative PretrainingEvaluationModel Merging
MP

Maura Pintor

University of Cagliari
Machine LearningAdversarial Machine LearningComputer Security
FK

Foutse Khomh

NSERC Arthur B. McDonald Fellow, CRC Tier 1, Canada CIFAR AI Chair, FRQ-IVADO Chair, Full Professor
Software engineeringMachine learning systems engineeringMining software repositoriesReverse