cross-architecture robustness comparison

Designs and executes experimental protocols and analysis pipelines to compare robustness properties across different model architectures and model families, producing comparable metrics and visualizations. This includes implementing standardized attacks and perturbations with controlled strength and normalization, class- and category-level breakdowns, and benchmarking procedures that highlight model-specific collapse or failure modes.

cross-architecturerobustnesscomparison

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.3
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Ensuring Robustness in ML-enabled Software Systems: A User Survey

Oct 21, 2025
HA
Hala Abdelkader
🏛️ Deakin University | RMIT University | Monash University

Machine learning systems face robustness challenges including silent failures, out-of-distribution (OOD) inputs, and adversarial attacks. To address these, this paper proposes ML-On-Rails—a standardized, engineering-oriented protocol. It introduces an HTTP status code–based model-software communication mechanism that enables structured, interpretable reporting of error types, prediction confidence, and decision rationales. The protocol integrates OOD detection, adversarial example identification, input validation, and local interpretability techniques, with design and validation rigorously informed by in-depth industrial practitioner surveys. Evaluation demonstrates that ML-On-Rails significantly improves fault observability and response consistency in ML systems, bridging a critical standardization gap in ML robustness assurance. It provides a practical, deployable engineering paradigm for building trustworthy ML systems.

Addressing silent failures and adversarial attacks in ML systemsEnhancing robustness against out-of-distribution data in ML softwareImproving transparency through standardized error reporting protocols

Existing software modeling datasets are often ad hoc constructions lacking rigorous quality assurance, leading to research findings that are difficult to reproduce, compare, and prone to bias. This work proposes the first benchmarking framework specifically designed for model-driven engineering, treating datasets themselves as first-class evaluation targets. By defining clear metrics for quality, representativeness, and task suitability, the framework establishes a unified platform that enables automated analysis of modeling datasets across multiple languages and formats. For the first time, this approach facilitates systematic evaluation of modeling datasets, substantially enhancing the reproducibility, fairness, and scientific rigor of research in the field.

benchmarkingdataset qualitymodel datasets

Reasonable Experiments in Model-Based Systems Engineering

Sep 12, 2025
JC
Johan Cederbladh
🏛️ Mälardalen University | Eindhoven University of Technology | Stellenbosch University | IT University of Copenhagen | University of Oslo | Universidade Federal Rural de Pernambuco | University of Antwerp

In model-based systems engineering, low experimental data reuse efficiency and excessive redundant experiments hinder digital engineering agility. To address this, this paper proposes a case-based reasoning (CBR)-driven experimental management framework that explicitly integrates domain knowledge. The framework features structured experimental metadata modeling, digital twin–enabled scenario semantic alignment, and an interpretable similarity assessment mechanism to intelligently determine whether historical experiments can be transferred to address new verification queries. Its key innovation lies in embedding domain knowledge explicitly into both the CBR retrieval and adaptation stages, thereby enabling trustworthy cross-operating-condition and cross-configuration experimental data reuse. Evaluated on an industrial-scale vehicle energy system design case, the framework reduces redundant experiments by 37% and shortens early verification cycles by 42% on average, significantly enhancing iterative efficiency in digital engineering and advancing intelligent experimental management.

Deciding if existing experiments can answer new engineering questionsIntelligently reusing experiment-related data to avoid redundant experimentsManaging experimental configuration metadata and results efficiently

This study addresses the limitations of existing SysML verification approaches, which are often tool-dependent and restricted to performance properties, lacking support for automated validation of behavioral and interface requirements. To overcome these shortcomings, this work proposes a tool-agnostic, automated verification workflow driven by SysML test cases, integrating UML Testing Profile and behavioral diagram constructs to enable unified validation of multidimensional attributes—including behavior, timing, and state responses. The methodology was developed through a mixed-methods research strategy combining literature review and stakeholder interviews, and its efficacy was empirically validated across two independent SysML toolchains. The approach not only transcends the constraints of conventional parametric methods but also enables automatic traceability of verification results back to the original model elements.

behavioral propertiesinterface propertiesmodel verification

Hot Scholars

BR

Bharat Runwal

IBM Research
Adversarial RobustnessSparsityGraph Neural Networks
OS

Olga Saukh

TU Graz / CSH Vienna
embedded intelligencemachine learningdeep learningsensing
RP

Rameswar Panda

Distinguished Engineer, IBM Research
Computer VisionMachine LearningNatural Language Processing
SK

Samir Khaki

University of Toronto
Artificial Intelligence
LB

Luca Benini

ETH Zürich, Università di Bologna
Integrated CircuitsComputer ArchitectureEmbedded SystemsVLSI